Model-based Multi-agent Mean-field Reinforcement Learning Based on Conditional Generation Adversarial Network Algorithm
Liubin Song, Daoxing Guo, Ao Wang, Dapeng Li, Feiyang Miao, Wenqi Sun · 2025
In the model-based multi-agent mean-field upper-confidence reinforcement learning ($\mathbf{M}^{\mathbf{3}}$-UCRL) framework, the large agent population and the environment’s complexity and variability introduce significant inaccuracies into the dynamics model’s next-state predictions.We propose a model-based multi-agent mean-field reinforcement learning based on conditional generative adversarial network (CGAN-M3RL) algorithm. The proposed algorithm utilized the adversarial training mechanism of conditional generative adversarial networks (CGAN) to establish an environmental dynamics model in multi-agent systems, which not only effectively improves the accuracy of the environmental model in predicting the next state but also provides an ample amount of sample data for subsequent policy learning, thereby enhancing the speed of policy convergence. Experimental evidence shows, with the $\mathrm{M}^{3}$-UCRL algorithm selected as a comparative baseline, that the CGAN-M3RL algorithm can effectively improve the accuracy of environmental dynamics model in predicting the next state.