Cooperation Pattern Exploration for Multi-Agent Reinforcement Learning
Chenran Zhao, Kecheng Peng, Xiaoqun Cao · 2022
Effective exploration is an important problem in reinforcement learning. It has developed well in the singleagent domain, but not so well in the multi-agent domain with more complex interactions. In this article, we provide a new approach to address the problem of efficient multi-agent exploration, which is called cooperation pattern exploration (CPE). CPE utilizes variational inference to predict cooperation patterns from the agent's observation-action trajectories, and trains the agent to distinguish and explore different cooperation patterns by minimizing the error from the true cooperation pattern. In addition, CPE applies the value decomposition framework to fuse the individual Q-values of agents during the centralized training stage. In order to enhance the effective exploration ability of agents, we integrate the novelty weight generated by random network distillation (RND) into the design of the value decomposition network to guide agents to explore more novel cooperation patterns. We test the performance of the designed algorithm on the SMAC platform. The experimental results show that CPE can effectively improve the cooperation ability of agents.