Dynamic Task Allocation for UAV Swarms in Maritime Rescue Scenarios Based on PG-MAPPO
Xiang Wu, Qingzhong Yan, Jiacun Wang, Yuhang Zhou, Qilong Huang, Changhui Jiang · IEEE Internet of Things Journal · 2025
The applications of unmanned swarms have become increasingly widespread, gradually transforming production processes and daily life. Task allocation, the top-level design for unmanned swarm missions, is pivotal to maximizing the efficiency of the entire swarm. However, traditional optimization methods and intelligent algorithms, including Reinforcement Learning (RL), often struggle to adapt to the complex and unpredictable situations in these tasks. To address this challenge, we propose a novel Multi-Agent Proximal Policy Optimization (MAPPO) algorithm combined with the population-based learning and Gaussian Mixture Model (GMM)-based adjustment mechanisms (PG-MAPPO). In PG-MAPPO, the population-based learning mechanism is integrated to enable agents with diverse exploration preferences to uncover optimal collaboration patterns among Unmanned Aerial Vehicles (UAVs), thereby enhancing cooperative efficiency. The GMM-based adjustment mechanism dynamically adjusts UAV formations for each agent, significantly improving the swarm’s flexibility and adaptability in rapidly changing environments. To demonstrate the effectiveness of PG-MAPPO, a maritime rescue simulation containing multiple complex and dynamic scenarios is conducted. Experimental results show that our algorithm achieves higher rescue success rate with faster convergence and greater stability than state-of-the-art Multi-Agent Reinforcement Learning (MARL) methods in all scenarios. Notably, the PG-MAPPO algorithm improves the rescue success rate by 31.6% compared to the best-performing baseline under challenging conditions.