Fast-QMIX: Accelerating Deep Multi-Agent Reinforcement Learning with Virtual Weighted Q-values

Boyang Yu, Zhaonian Cai, Jingbo He · 2021 2nd International Conference on Electronics, Communications and Information Technology (CECIT) · 2021

Cooperation between agents in a multi-agent system (MAS) is prevalent in real or virtual environments. In a fully cooperative environment, each agent learns its own policy by using the overall reward. QMIX adds an additional parametric network to replace the linear summation of values in Value-Decomposition Networks (VDNs), aiming to learn more complex relationships between agents. However, QMIX can not determine the best value for each agent during the process, resulting in slower convergence, instability of training, and poor performance. In this paper, we further dynamically assign virtual weighted Q-values based on QMIX with an additional network to further utilize the global information as an auxiliary guide and expand the ability to explore while also expand the robustness of the algorithm. Experiments show that our method not only outperforms original QMIX in different scenarios of the StarCraft Multi-Agent Challenge (SMAC) but also converges faster and more stable with the addition of calibration.

Read the paper · More papers on PaperTik