Learning maximum entropy policies with QMIX in cooperative MARL
Fenggen Guo, Zizhao Wu · 2022 IEEE 2nd International Conference on Electronic Technology, Communication and Information (ICETCI) · 2022
Model-free deep reinforcement learning have been successfully applied to a range of challenging sequential decision making tasks. However, most tasks in the real world consist of multi-agent and require cooperative behaviors among agents. Quite a few methods such as QMIX have been proposed to address the credit assignment problem and learn cooperative policies in MARL so far. However, such methods suffer from lack of deep and adequate explorations, causing miss of the optimal policies. We propose a novel method QMIX-ME for MARL in this paper that learns policies with maximum entropy for exploration inspired by SAC. For the compatibility of applying policies with maximum entropy in MARL, we make modifications as followed. First, we set independent entropy temperatures for each agent and make approximations for total state values. Besides, we adopt self-attention mechanism in individual critic networks to capture relations and avail of information of other agents. We evaluate our method on StarCraft Multi-Agent Challenge (SMAC) and the experiments results show great progress of our method.