Collaborative decision modeling of multi-agent based on BC-QMIX for air mission wargaming

Ze Wang, Ni Li, Guanghong Gong · Advances in Complex Systems · 2025

The existing methods for air mission wargaming autonomous decision-making, such as optimization theory-based and expert systems, often suffer from insufficient real-time performance and high modeling workload. The application of multi-agent reinforcement learning (RL) in autonomous decision-making for air mission wargaming has received extensive attention. However, most existing RL approaches require online interaction, leading to low efficiency of sample collection and network training. This paper proposes BC-QMIX, the network structure of BC-QMIX employs supervised learning method to train a behavior cloning network for each sub-agent on the basis of the QMIX network, which provides a basis for action selection. Therefore, BC-QMIX alleviates the extrapolation error in the offline training of QMIX. Furthermore, when offline pre-training is conducted by utilizing domain knowledge-based samples, it accelerates the online training and convergence of network. Through experiments conducted in the Multi Drones Monitoring environment and two collaborative air mission wargaming scenarios, BC-QMIX demonstrates a significant reduction in extrapolation error and outperforms QMIX, MADDPG, and MATD3 in offline training. Specifically, it achieves a 10.1% improvement in average winning rate over QMIX, with this enhancement rising to 47.9% when domain knowledge is incorporated into the training process. This validation demonstrates the feasibility and advantage of constructing a collaborative autonomous decision-making model using BC-QMIX for air mission wargaming scenarios.

Read the paper · More papers on PaperTik