A Multi-Agent Q-Learning with Value Function Approximation Based on Single-leader Multi-followers Stackelberg Game

Chunxi Zhu, Wenwu Yu, He Wang · 2023

Although multi-agent reinforcement learning (MARL) has made significant progress in dealing with complex tasks, the hypothesis that agents act simultaneously still limits the applicability of MARL in many real-world problems. In this work, we relax such hypothesis by proposing a single-leader multi-followers Stackelberg game model. In this model, the leader considers the policy of the followers, and the followers make the best response based on the leader's action. By combining the single-leader multi-followers Stackelberg game model with Q-learning, we propose a Q-learning algorithm based on the Stackelberg game. We test a series of games, including Lumberjack and Predator Prey, which are challenging for existing MARL algorithms. Experimental results show that our method achieves competitive advantages in terms of performance and convergence speed.

Read the paper · More papers on PaperTik