Heterogeneous Policy Network Reinforcement Learning for UAV Swarm Confrontation
Aocheng Su, Fengxuan Hou, Yiguang Hong · 2024
UAV swarms confrontation is becoming increasingly important in the field of military security, however, learning effective cooperative strategies for heterogeneous UAV swarms confrontation remains a major challenge. In this paper, we address the heterogeneous UAV swarms confrontation problem by developing the UAV swarms collaborative state space reconstruction mechanism. This mechanism reconstructs the UAV's own state space by using the relative state information to depict the collaborative relationship between the UAV swarms. Then, within the Multi-Agent Proximal Policy Optimization (MAPPO) deep reinforcement learning framework, we propose a heterogeneous policy network local parameter sharing MAPPO algorithm. This algorithm addresses the differences in heterogeneous agent state spaces and action decision spaces by designing heterogeneous policy networks for different types of agents and optimizing the policy network parameters using the local parameter sharing mechanism. Finally, with a simulation environment for reconnaissance-attack heterogeneous UAV swarms, experiments comparing the proposed method with existing multi-agent reinforcement learning algorithms demonstrate the effectiveness of our approach.