An Algorithm Combining Hidden States for Monotonic Value Function Factorisation

Ershen Wang, Xiaotong Wu, Chen Hong, Jihao Chen · 2023

By using the communication and shared cooperation information between agents to infer the observation states of other agents, partially observable problems in multi-agent systems can be solved by centralized training decentralized execution (CTDE). However, even in CTDE, agents may still get stuck in local optimal solutions. In order to solve the problem, we design a novel multi-agent reinforcement learning (MARL) algorithm, hidden state based QMIX (HSQMIX), which uses the hidden state generated by recurrent neural networks (RNNs) to replace the true state of the agent as the input of the mixing network, and introduces transformer to process the hidden state. Besides, the priority experience replay and temporal difference learning are elaborately integrated into the algorithm. We also evaluate HSQMIX in the StarCraft multi-agent challenge (SMAC) map and compare it with other popular MARL algorithms. Experimental results show that HSQMIX outperforms other algorithms by a large margin. Our work will provide new insights into the multi-agent cooperative game.

Read the paper · More papers on PaperTik