Multi-Attacker Multi-Defender Target Guarding Game Using Hierarchical Reinforcement Learning
Shaoqian Dong, Yi C. Huang · 2025
This study proposes a Hierarchical Reinforcement Learning (HRL) framework for multi-agent target guarding in dynamic environments with respectively coordinated attackers and defenders. The framework decomposes defender actions into patrol, pursuit, and encirclement subtasks, with a high-level multi-head attention mechanism dynamically allocating subtasks based on global observations. Defenders trained with Multi-Agent Proximal Policy Optimization (MAPPO) execute a low-level policy to engage subtasks. Customized reward functions promote collision avoidance, target protection, and coordination: patrol rewards optimize circular surveillance, pursuit rewards minimize distances to attackers, and encirclement rewards enhance cooperation. Attackers employ evasion tactics to breach defenses. Simulations in 2D environments demonstrate effective subtask transitions and coordinated interceptions, validating the framework's robustness against nonstationary interactions.