An Improved WQMIX with Dual-Network and Dynamic Weights for Cooperative Multi-Agent Reinforcement Learning

Zhehan Xing, Jie Luo, Hao Zhao · 2024

In the domain of cooperative multi-agent reinforcement learning (MARL), QMIX is a key algorithm that facilitates effective collaboration among multiple agents by enforcing monotonicity constraints on the joint Q-values. However, this assumption of monotonicity may limit the algorithm's capacity to explore more complex strategies. As an extension of QMIX, WQMIX attempts to relax these constraints by incorporating a weight scheme, aimed at enhancing the algorithm's flexibility and efficiency when handling complex multi-agent interactions. Nevertheless, its weight design is overly simplistic and fails to fully leverage potential performance. This paper introduces an improved version of WQMIX, termed Dual-Network and Dynamic Weights QMIX (DDWQMIX), which employs a dynamic weights strategy to more flexibly adapt to changes in the environment and optimize the learning process. Additionally, DDWQMIX integrates a dual-network structure, updating targets based on the smaller Q-values produced by the two networks, thus allowing for a more accurate estimation of joint Q-values and ensuring that optimal actions are not missed. Experiments conducted in the starcraft multi-agent challenge (SMAC) environment demonstrate that DDWQMIX outperforms traditional VDN, QMIX and WQMIX in multi-agent cooperative tasks, particularly in more complex scenarios.

Read the paper · More papers on PaperTik