Multi-Robot Flocking Control Using Multi-Agent Twin Delayed Deep Deterministic Policy Gradient

Mario Salama Youssef, Nouran Adel Hassan, Ayman El-Badawy · 2022 19th International Conference on Electrical Engineering, Computing Science and Automatic Control (CCE) · 2022

This paper proposes the use of Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) to resolve the issues accompanied with Multi-Agent Deep Deterministic Policy Gradient (MADDPG) such as the overestimation bias of a centralized critic when used to solve the multi-robot flocking control problem. A temporal difference error prioritized replay buffer is used alongside to achieve the most efficient learning process. In order to allow the robots to maintain a flock throughout a complex environment, an appropriate reward function was constructed, taking into consideration the following parameters: reaching the goal, maintaining a distance between the agents to ensure a stable connection while avoiding collision between one another, obstacle avoidance, and moving in a specific formation. Results of both algorithms are then compared to test the performance of the MATD3.

Read the paper · More papers on PaperTik