The End-to-End Cooperative Hunting Method of UAVs Driven by Layered Rewards
Suqin Wu, Shengyang Liu, Haolong Feng, Ting Song, Fei Han · 2025
For tackling the time-varying and cooperative hunting mission challenge of unmanned aerial vehicles, an end-to-end cooperative hunting method driven by the layered reward is proposed. A centralized training distributed execution architecture named MADDPG-Attention of the multi-agent reinforcement learning is adapted to interact with the environment to generate the decision model of each agent. During the interaction with the environment, a layered reward mechanism is put forward to drive each agent to train its own critic network. The reasonable two layer reward ensures the quick convergence of the algorithm and a good hunting performance. Finally, a 3vs1 hunting experiment of UAVs is conducted and the results effectively demonstrate the advantages of the proposed hunting method interms of stability and accuracy compared to the benchmark algorithm.