The End-to-End Cooperative Hunting Method of UAVs Driven by Layered Rewards

Suqin Wu, Shengyang Liu, Haolong Feng, Ting Song, Fei Han · 2025

For tackling the time-varying and cooperative hunting mission challenge of unmanned aerial vehicles, an end-to-end cooperative hunting method driven by the layered reward is proposed. A centralized training distributed execution architecture named MADDPG-Attention of the multi-agent reinforcement learning is adapted to interact with the environment to generate the decision model of each agent. During the interaction with the environment, a layered reward mechanism is put forward to drive each agent to train its own critic network. The reasonable two layer reward ensures the quick convergence of the algorithm and a good hunting performance. Finally, a 3vs1 hunting experiment of UAVs is conducted and the results effectively demonstrate the advantages of the proposed hunting method interms of stability and accuracy compared to the benchmark algorithm.

Read the paper · More papers on PaperTik