Application of Self-Play Deep Reinforcement Learning in Task Planning for UAV Swarms under Capacity Constraints

Jiaxi Chen, Zilong Chen, Pengcheng Wen, Binqi Chen, Wenjun Yang · 2025

This paper develops an Adaptive Staircase-based Policy Space Response Oracle (ASP) architecture that uses a self-play deep reinforcement learning (DRL) method to solve the problem of planning tasks for UAV swarms in urban logistics and post-disaster rescue situations when there isn't enough time and resources. We show that UAV swarm path planning may be optimized across different scales by combining a two-player meta-game structure with adversarial instance creation and developing a staircase curriculum learning mechanism. The framework lets you learn from small amounts of pre-training (10 to 100 inspection points) and apply that knowledge to larger applications (200 to 1000 inspection points). The results of the experiments indicate that the ASP-based method works better than traditional Policy Optimization for Multi-Agent (POMO) algorithms. It cuts the average solution time by 41% and the total flight distance by 55.6% in tasks with a thousand nodes. This paper offers a systematic engineering approach for dynamic task planning in UAV swarms, overcoming the problem of balancing exploration and exploitation in complicated state spaces and proving that it can be used in a variety of situations.

Read the paper · More papers on PaperTik