中国机械工程 ›› 2026, Vol. 37 ›› Issue (8): 2017-2028.DOI: 10.3969/j.issn.1004-132X.2026.08.021
• 智能制造 • 上一篇
收稿日期:2025-08-07
出版日期:2026-08-25
发布日期:2026-09-17
通讯作者:
郑鹏
作者简介:肖世昌,男,1987年生,副教授、博士。研究方向为港口调度优化/智能制造系统建模与调度优化。E-mail: scxiao@shmtu.edu.cn
XIAO Shichang1(
), LIN Xuan1, ZHENG Peng1(
), WANG Jinfeng2
Received:2025-08-07
Online:2026-08-25
Published:2026-09-17
Contact:
ZHENG Peng
摘要:
为了提高自动化集装箱码头的运营效率和智能决策能力,针对水平运输区域中双向行驶的多自动导引车(AGV)防冲突路径规划问题展开研究。首先针对码头水平运输区域特征构建网格地图,并将AGV防冲突路径规划问题建模为以最小化所有任务最大完工时间为目标的数学规划模型。随后设计了近端策略优化(PPO)算法,搭建了面向双向引导车道的工作环境模型,设计了多AGV作业的动作空间和状态空间,通过引入防绕远策略提高算法搜索能力。最后将所提算法与商业求解器Gurobi、A*算法及遗传算法进行性能对比,仿真实验结果表明,所提算法针对大规模问题具有更强的性能稳定性和收敛能力。
中图分类号:
肖世昌, 林轩, 郑鹏, 王金凤. 基于近端策略优化算法的自动化集装箱码头自动导引车防冲突路径规划[J]. 中国机械工程, 2026, 37(8): 2017-2028.
XIAO Shichang, LIN Xuan, ZHENG Peng, WANG Jinfeng. Anti-conflict Path Planning for AGVs in the Automated Container Terminals Based on Proximal Policy Optimization Algorithm[J]. China Mechanical Engineering, 2026, 37(8): 2017-2028.
| 参数 | 含义 |
|---|---|
| AGV从当前位置向上运动是否会超过边界 | |
| AGV从当前位置向下运动是否会超过边界 | |
| AGV从当前位置向左运动是否会超过边界 | |
| AGV从当前位置向右运动是否会超过边界 | |
| AGV当前所处位置的横坐标 | |
| AGV当前所处位置的纵坐标 | |
| AGV当前位置与目标位置的相对横坐标 | |
| AGV当前位置与目标位置的相对纵坐标 | |
| 当前AGV与附近最近AGV之间的距离 |
表1 动作空间参数
Tab.1 Action space parameters
| 参数 | 含义 |
|---|---|
| AGV从当前位置向上运动是否会超过边界 | |
| AGV从当前位置向下运动是否会超过边界 | |
| AGV从当前位置向左运动是否会超过边界 | |
| AGV从当前位置向右运动是否会超过边界 | |
| AGV当前所处位置的横坐标 | |
| AGV当前所处位置的纵坐标 | |
| AGV当前位置与目标位置的相对横坐标 | |
| AGV当前位置与目标位置的相对纵坐标 | |
| 当前AGV与附近最近AGV之间的距离 |
| PPO算法步骤 |
|---|
Input: Output: 最大完工时间 1. 2. 3. 4. repeat 5. 根据state和 6. 7. 8. 根据 9. state = 10. 根据state计算奖励reward 11. list = 12. buffer = buffer + list ∥将list记录到经验数组buffer中 13. if buffer_length mod 2^9 == 0 then 14. 根据buffer,B,R,使用PPO算法更新神经网络参数 15. 16. until [ 17. 18. 19. return |
表2 基于PPO算法的AGV路径规划算法步骤
Tab.2 Implementation steps of AGV path planning algorithm based on PPO
| PPO算法步骤 |
|---|
Input: Output: 最大完工时间 1. 2. 3. 4. repeat 5. 根据state和 6. 7. 8. 根据 9. state = 10. 根据state计算奖励reward 11. list = 12. buffer = buffer + list ∥将list记录到经验数组buffer中 13. if buffer_length mod 2^9 == 0 then 14. 根据buffer,B,R,使用PPO算法更新神经网络参数 15. 16. until [ 17. 18. 19. return |
| 参数 | 取值 |
|---|---|
| 学习率 | 0.000 007 |
| 折扣因子 | 0.88 |
| 裁剪参数 | 0.2 |
| 批次大小(batch size) | 16 |
| 单轮更新的采样步数(sample step) | 512 |
| 数据复用次数(reuse times) | 2 |
| 限制新旧策略总体差异的系数 | 0.025 |
| GAE调整方差与偏差的系数 | 0.98 |
表3 PPO算法的相关超参数
Tab.3 Hyperparameters of PPO algorithm
| 参数 | 取值 |
|---|---|
| 学习率 | 0.000 007 |
| 折扣因子 | 0.88 |
| 裁剪参数 | 0.2 |
| 批次大小(batch size) | 16 |
| 单轮更新的采样步数(sample step) | 512 |
| 数据复用次数(reuse times) | 2 |
| 限制新旧策略总体差异的系数 | 0.025 |
| GAE调整方差与偏差的系数 | 0.98 |
| 地图大小 | AGV个数 | 最大完工时间(Makespan) | 求解时间/s | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Gurobi | A* | GA | 改进PPO | Gurobi | A* | GA | 改进PPO | ||
| 10×10 | 3 | 15 | 15 | 15 | 15 | 4.33 | 0.01 | 0.50 | 0.37 |
| 5 | 17 | 19 | 19 | 16 | 43.76 | 0.02 | 1.85 | 0.97 | |
| 10 | 18 | 18 | 18 | 21 | 478.96 | 0.03 | 3.67 | 1.91 | |
| 20×20 | 5 | 31 | 31 | 31 | 0.08 | 4.14 | 0.62 | ||
| 10 | 37 | 35 | 36 | 0.27 | 8.15 | 1.37 | |||
| 15 | 34 | 34 | 34 | 0.34 | 11.61 | 2.06 | |||
| 30×30 | 20 | 47 | 49 | 49 | 4.39 | 26.38 | 2.11 | ||
| 25 | 43 | 43 | 43 | 7.56 | 25.57 | 3.31 | |||
| 30 | 45 | 45 | 46 | 13.52 | 36.54 | 5.06 | |||
表4 三种算法与改进PPO算法在不同规模下的性能对比
Tab.4 Performance comparison of three algorithms and the improved PPO under different scales
| 地图大小 | AGV个数 | 最大完工时间(Makespan) | 求解时间/s | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Gurobi | A* | GA | 改进PPO | Gurobi | A* | GA | 改进PPO | ||
| 10×10 | 3 | 15 | 15 | 15 | 15 | 4.33 | 0.01 | 0.50 | 0.37 |
| 5 | 17 | 19 | 19 | 16 | 43.76 | 0.02 | 1.85 | 0.97 | |
| 10 | 18 | 18 | 18 | 21 | 478.96 | 0.03 | 3.67 | 1.91 | |
| 20×20 | 5 | 31 | 31 | 31 | 0.08 | 4.14 | 0.62 | ||
| 10 | 37 | 35 | 36 | 0.27 | 8.15 | 1.37 | |||
| 15 | 34 | 34 | 34 | 0.34 | 11.61 | 2.06 | |||
| 30×30 | 20 | 47 | 49 | 49 | 4.39 | 26.38 | 2.11 | ||
| 25 | 43 | 43 | 43 | 7.56 | 25.57 | 3.31 | |||
| 30 | 45 | 45 | 46 | 13.52 | 36.54 | 5.06 | |||
| AGV编号 | 起始位置坐标 | 终点位置坐标 | 任务类型 |
|---|---|---|---|
| 1 | (0,1) | (10,4) | 前往岸桥装箱 |
| 2 | (0,3) | (10,7) | 前往岸桥装箱 |
| 3 | (10,3) | (0,6) | 前往堆场卸箱 |
| 4 | (10,6) | (0,8) | 前往堆场卸箱 |
| 5 | (10,9) | (0,2) | 前往堆场卸箱 |
表5 5辆AGV起始位置与终点位置
Tab.5 Starting and ending positions of 5 AGVs
| AGV编号 | 起始位置坐标 | 终点位置坐标 | 任务类型 |
|---|---|---|---|
| 1 | (0,1) | (10,4) | 前往岸桥装箱 |
| 2 | (0,3) | (10,7) | 前往岸桥装箱 |
| 3 | (10,3) | (0,6) | 前往堆场卸箱 |
| 4 | (10,6) | (0,8) | 前往堆场卸箱 |
| 5 | (10,9) | (0,2) | 前往堆场卸箱 |
| 地图大小 | AGV个数 | PPO算法是否考虑 防绕远机制 | 训练轮数 | 最大完工时间 (Makespan) | AGV完工时间之和 | AGV路径步数之和 |
|---|---|---|---|---|---|---|
| 10×10 | 5 | 否 | 300 | 18 | 80 | 79 |
| 是 | 300 | 16 | 69 | 73 | ||
| 10×10 | 10 | 否 | 300 | 23 | 195 | 185 |
| 是 | 300 | 21 | 168 | 167 |
表6 两种PPO算法求解双向车道上的AGV路径规划结果对比
Tab.6 Comparison of path planning results for AGVs on bidirectional lanes using two PPO algorithms
| 地图大小 | AGV个数 | PPO算法是否考虑 防绕远机制 | 训练轮数 | 最大完工时间 (Makespan) | AGV完工时间之和 | AGV路径步数之和 |
|---|---|---|---|---|---|---|
| 10×10 | 5 | 否 | 300 | 18 | 80 | 79 |
| 是 | 300 | 16 | 69 | 73 | ||
| 10×10 | 10 | 否 | 300 | 23 | 195 | 185 |
| 是 | 300 | 21 | 168 | 167 |
| [1] | ROY D, de KOSTER R, BEKKER R. Modeling and Design of Container Terminal Operations[J]. Operations Research, 2020, 68(3): 686-715. |
| [2] | YUAN Ruiping, DONG Tingting, LI Juntao. Research on the Collision-free Path Planning of Multi-AGVs System Based on Improved A* Algorithm[J]. American Journal of Operations Research, 2016, 6(6): 442-449. |
| [3] | YANG Yongsheng, ZHONG Meisu, DESSOUKY Y, et al. An Integrated Scheduling Method for AGV Routing in Automated Container Terminals[J]. Computers & Industrial Engineering, 2018, 126: 482-493. |
| [4] | Min LÜ, GAO Tong, ZHANG Nian. Research of AGV Scheduling and Path Planning of Automatic Transport System[J]. International Journal of Control and Automation, 2016, 9(4): 1-12. |
| [5] | KIM C W, TANCHOCOJ J M A. Operational Control of a Bidirectional Automated Guided Vehicle System[J]. International Journal of Production Research, 1993, 31(9): 2123-2138. |
| [6] | WANG Zehao, ZENG Qingcheng. A Branch-and-bound Approach for AGV Dispatching and Routing Problems in Automated Container Terminals[J]. Computers & Industrial Engineering, 2022, 166: 107968. |
| [7] | 高一鹭, 胡志华. 基于时空网络的自动化集装箱码头自动化导引车路径规划[J]. 计算机应用, 2020, 40(7): 2155-2163. |
| GAO Yilu, HU Zhihua. Path Planning for Automated Guided Vehicles Based on Tempo-spatial Network at Automated Container Terminal[J]. Journal of Computer Applications, 2020, 40(7): 2155-2163. | |
| [8] | SHEN Lixin, WANG Yaodong, LIU Kunpeng, et al. Synergistic Path Planning of Multi-UAVs for Air Pollution Detection of Ships in Ports[J]. Transportation Research Part E: Logistics and Transportation Review, 2020, 144: 102128. |
| [9] | 梁承姬, 沈珊珊, 胡文辉. 基于路段时间窗考虑备选路径的AGV路径规划[J]. 工程设计学报, 2018, 25(2): 200-208. |
| LIANG Chengji, SHEN Shanshan, HU Wenhui. AGV Path Planning Considering Alternative Paths Based on Time Window of Road Section[J]. Chinese Journal of Engineering Design, 2018, 25(2): 200-208. | |
| [10] | HU Yejun, DONG Liangcai, XU Lei. Multi-AGV Dispatching and Routing Problem Based on a Three-stage Decomposition Method[J]. Mathematical Biosciences and Engineering, 2020, 17(5): 5150-5172. |
| [11] | XU Yixiang, QI Liang, LUAN Wenjing, et al. Load-in-load-out AGV Route Planning in Automatic Container Terminal[J]. IEEE Access, 2020, 8: 157081-157088. |
| [12] | 姜辰凯, 李智, 盘书宝, 等. 基于改进Dijkstra算法的AGVs无碰撞路径规划[J]. 计算机科学, 2020, 47(8): 272-277. |
| JIANG Chenkai, LI Zhi, PAN Shubao, et al. Collision-free Path Planning of AGVs Based on Improved Dijkstra Algorithm[J]. Computer Science, 2020, 47(8): 272-277. | |
| [13] | TANG Gang, TANG Congqiang, CLARAMUNT C, et al. Geometric A-star Algorithm: an Improved A-star Algorithm for AGV Path Planning in a Port Environment[J]. IEEE Access, 2021, 9: 59196-59210. |
| [14] | 丁一, 袁浩, 方怀瑾, 等. 考虑冲突规避的自动化集装箱码头AGV优化调度方法[J]. 交通信息与安全, 2022, 40(3): 96-107. |
| DING Yi, YUAN Hao, FANG Huaijin, et al. An Optimal Scheduling Method of AGVs at Automated Container Terminal Considering Conflict Avoidance[J]. Journal of Transport Information and Safety, 2022, 40(3): 96-107. | |
| [15] | ZHONG Meisu, YANG Yongsheng, DESSOUKY Y, et al. Multi-AGV Scheduling for Conflict-free Path Planning in Automated Container Terminals[J]. Computers & Industrial Engineering, 2020, 142: 106371. |
| [16] | WU Maopu, GAO Jian, LI Le, et al. Control Optimisation of Automated Guided Vehicles in Container Terminal Based on Petri Network and Dynamic Path Planning[J]. Computers and Electrical Engineering, 2022, 104: 108471. |
| [17] | CHEN Tianjian, SUN Yuan, DAI Wei, et al. On the Shortest and Conflict-free Path Planning of Multi-AGV System Based on Dijkstra Algorithm and the Dynamic Time-window Method[J]. Advanced Materials Research, 2013, 645: 267-271. |
| [18] | LIU Wenqian, ZHU Xiaoning, WANG Li, et al. Multiple Equipment Scheduling and AGV Trajectory Generation in U-shaped Sea-rail Intermodal Automated Container Terminal[J]. Measurement, 2023, 206: 112262. |
| [19] | CAO Menglong, PENG Zhang. Research on Loading and Unloading Path Optimization for AGV at Automatic Container Terminal Based on Improved Particle Swarm Algorithm[C]∥2020 4th Annual International Conference on Data Science and Business Analytics (ICDSBA). IEEE, 2020: 1-5. |
| [20] | LU Tong, SUN Zhaohui, QIU Siqi, et al. Time Window Based Genetic Algorithm for Multi-AGVs Conflict-free Path Planning in Automated Container Terminals[C]∥2021 IEEE International Conference on Industrial Engineering and Engineering Management (IEEM). IEEE, 2021: 603-607. |
| [21] | 岳春擂, 黄俊, 邓乐乐. 改进蚁群算法在AGV路径规划上的研究[J]. 计算机工程与设计, 2022, 43(9): 2533-2541. |
| YUE Chunlei, HUANG Jun, DENG Lele. Research on Improved Ant Colony Algorithm in AGV Path Planning[J]. Computer Engineering and Design, 2022, 43(9): 2533-2541. | |
| [22] | 荀燕琴. 基于群体智能优化的AGV路径规划算法研究[D]. 长春: 吉林大学, 2017. |
| XUN Yanqin. Research on AGV Path Planning Algorithm Based on Group Intelligence Optimization[D]. Changchun: Jilin University, 2017. | |
| [23] | CRUZ D L, YU Wen. Path Planning of Multi-agent Systems in Unknown Environment with Neural Kernel Smoothing and Reinforcement Learning[J]. Neurocomputing, 2017, 233: 34-42. |
| [24] | POPPER J, YFANTIS V, RUSKOWSKI M. Simultaneous Production and AGV Scheduling Using Multi-agent Deep Reinforcement Learning[J]. Procedia CIRP, 2021, 104: 1523-1528. |
| [25] | ÇETINKAYA M. Multi-agent Path Planning Using Deep Reinforcement Learning[J]. ArXiv Preprint ArXiv: , 2021. |
| [26] | CHEN Xinqiang, LIU Shuhao, ZHAO Jiansen, et al. Autonomous Port Management Based AGV Path Planning and Optimization via an Ensemble Reinforcement Learning Framework[J]. Ocean & Coastal Management, 2024, 251: 107087. |
| [27] | WANG Tingzhong, ZHANG Binbin, ZHANG Mengyan, et al. Multi-UAV Collaborative Path Planning Method Based on Attention Mechanism[J]. Mathematical Problems in Engineering, 2021, 2021: 6964875. |
| [28] | RIEDMILLER M, HAFNER R, LAMPE T, et al. Learning by Playing—Solving Sparse Reward Tasks from Scratch[C]∥Proceedings of the 35th International Conference on Machine Learning (ICML). Stockholmsmässan, 2018: 4344-4353. |
| [29] | 蔡泽, 胡耀光, 闻敬谦, 等. 复杂动态环境下基于深度强化学习的AGV避障方法[J]. 计算机集成制造系统, 2023(1): 236-245. |
| CAI Ze, HU Yaoguang, WEN Jingqian, et al. Collision Avoidance for AGV Based on Deep Reinforcement Learning in Complex Dynamic Environment[J]. Computer Integrated Manufacturing Systems, 2023(1): 236-245. | |
| [30] | SARTORETTI G, KERR J, SHI Yunfei, et al. PRIMAL: Pathfinding via Reinforcement and Imitation Multi-agent Learning[J]. IEEE Robotics and Automation Letters, 2019, 4(3): 2378-2385. |
| [31] | SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal Policy Optimization Algorithms[PP/OL].V2.arXiv(2017-08-28).. |
| [32] | CAO Yu, YANG Ang, LIU Yang, et al. AGV Dispatching and Bidirectional Conflict-free Routing Problem in Automated Container Terminal[J]. Computers & Industrial Engineering, 2023, 184: 109611. |
| [1] | 武星, 吕鹏, 沙金龙, 李杨志, 孟昭旭. 基于改进多目标粒子群优化的移动机器人对接轨迹规划[J]. 中国机械工程, 2026, 37(8): 1999-2008. |
| [2] | 陈亚洲, 王国荣, 董学成, 唐洋. 基于变径轮廓优化的控压节流阀线性设计[J]. 中国机械工程, 2026, 37(7): 1595-1602. |
| [3] | 董佳祥, 赵学智, 胡希平, 刘铨权. 一种软轴拉扭协同驱动的新型刚柔混联连续体机器人机构设计与实验验证[J]. 中国机械工程, 2026, 37(5): 1193-1198. |
| [4] | 曹轶然, 高磊, 吴孟丽, 王旭浩, 彭聪, 郭志永, 梁瑶. 融合预瞄机制的移动机器人预设性能视觉伺服控制策略[J]. 中国机械工程, 2026, 37(5): 1218-1225. |
| [5] | 肖伟, 张聪, 陈绪兵. 基于贝叶斯优化时间卷积网络的工业机器人能耗预测[J]. 中国机械工程, 2026, 37(4): 831-836. |
| [6] | 孙悦, 黄辉, 尹方辰. 高负载动态工况下工业机器人的能耗预测[J]. 中国机械工程, 2026, 37(4): 939-947. |
| [7] | 郭万金, 利乾辉, 田玉祥, 曹雏清, 赵立军, 徐明坤, 刘孝恒, 侯旭栋. 动力学模态解耦的工业机器人加工颤振规避方法[J]. 中国机械工程, 2026, 37(3): 571-585. |
| [8] | 孙立峰, 张永顺, 鲁天宇, 刘振虎, 孙丽敏. 电磁球型关节全悬浮转子抑振油腔廓形优化[J]. 中国机械工程, 2026, 37(2): 374-382. |
| [9] | 赵丁选, 郭瑞, 王硕, 闫长长, 王子鹤, 张天赐. 复杂地形环境下无人步履式挖掘机的车身姿态规划方法[J]. 中国机械工程, 2026, 37(1): 233-242. |
| [10] | 郭万金, 田玉祥, 利乾辉, 曹雏清, 赵立军, 徐明坤, 刘孝恒, 侯旭栋. 未知环境下机器人打磨自适应变阻抗恒力控制[J]. 中国机械工程, 2026, 37(1): 92-104. |
| [11] | 妥吉英, 徐笑南, 李俊, 张玉琛, 黄安, 胡都, 刘梓林. 一种基于改进SAC算法的六轴机械臂路径规划[J]. 中国机械工程, 2025, 36(12): 2986-2992. |
| [12] | 张雷, 杨聪楠, 李崴一, 赵一洁, 王晓聪. 变高度双轮足平台自适应平衡控制算法设计[J]. 中国机械工程, 2025, 36(12): 2920-2926. |
| [13] | 倪涛, 赵亚辉, 赵泽仁, 杨凯强. 6-UPRU并联机器人动力学建模及基本动力学参数确定[J]. 中国机械工程, 2025, 36(12): 2911-2919. |
| [14] | 杨明星, 沈佳乐, 高鹏, 张兴, 王俊翔. 连续体机器人设计与导向路径损失补偿策略[J]. 中国机械工程, 2025, 36(12): 2820-2828. |
| [15] | 梁海平, 卢耀安, 连伟嘉, 王成勇. 考虑冗余自由度的六轴机器人光顺运动路径规划方法[J]. 中国机械工程, 2025, 36(11): 2652-2657. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||