Homogeneous Decision Networks for Multi-Agent Formation Control via Distributed Reinforcement Learning
Yonghao Li, S. Kevin Zhou, Muzhen Lyu, Dongze Li, Zhuo Zou, Lizheng Liu · Procedia Computer Science · 2025
Aiming at the problem of balancing formation maintenance, obstacle avoidance and dynamic navigation in multi-robot cooperative formation control in complex environments, this study proposes a homogeneous decision network framework based on distributed reinforcement learning. The framework adopts parameter-shared homogeneous policy networks, where all robots have identical decision networks, and decision differences only arise from local observation states such as relative positions, obstacle perception, and target directions. Improved PPO and A3C algorithms are used to handle multiple tasks, with the gradient aggregation mechanism of A3C to synchronize policy parameters, and a dual-critic network is constructed to dynamically balance the weights of individual and global rewards. A dynamic Leader-Follower architecture is adopted to elect Leaders based on proximity to target points and design differentiated reward functions, with all logics integrated into a single homogeneous network, enabling agents to autonomously assume different roles and adjust formations according to real-time states. Simulation experiments verify the effectiveness of the algorithm through a three-stage progressive training scheme.