A Multi-UAV Monitoring and Search Strategy Based on Multi-Agent Reinforcement Learning

Shuyi Guo, Xuetao Zhang, Hanzhang Wang, Gang Chen Sun, Yisha Liu, Yan Zhuang · 2024

Multiple UAVs have been widely used for targets search and monitoring. Nonetheless, in practice, this problem is challenging as targets usually move randomly, but the trajectories of the targets cannot be predicted in advance by UAV swarms. Moreover, the efficiency of exploration, which indirectly impacts monitoring, is low when the number of targets in the environment is unknown. A Search Extended Deep Deterministic Policy Gradient (SEDDPG) method is proposed by utilizing distributed training and information sharing among multiple agents. Specifically, a f ramework o f d ecentralized p artially observable Markov decision processes is established to express this multiobjective optimization problem. Through distributed training and information sharing among multiple agents, a Search Extended Deep Deterministic Policy Gradient (SEDDPG) method is proposed to synchronize exploration and monitoring tasks. This helps multiple UAVs to monitor targets in the environment as extensively as possible. Simulation results demonstrate that the framework is converged, and outperforms in terms of monitoring ability and search efficiency.

Read the paper · More papers on PaperTik