Learning Cooperative Multi-Agent Policies with Multi-Channel Reward Curriculum Based Q-Learning

Jayant Kumar Singh, Jing Lun Zhou, Baltasar Beferull‐Lozano, Ilya Tyapin · IECON 2022 – 48th Annual Conference of the IEEE Industrial Electronics Society · 2022

Multi-Agent Reinforcement Learning (MARL) algorithms based on the Centralised Training Decentralised Execution (CTDE) approach have seen a great deal of interest in recent years. Most of the recent works focus on a specific class of environments built on the StartCraft Multi-Agent Challenge (SMAC) environment suite. However, experiments with the PettingZoo multi-particle environments show poor performance in an interesting subset of tasks. In this paper, the nature of these environments and the reward structures are analyzed. It shows that poor performance in these tasks is not due to a lack of representation power in the individual Q function or mixing functions, but rather a result of convergence to a suboptimal equilibrium of dual channel rewards and the issue of agent-reward decoupling that can be a common problem for many MARL environments. We present reward curriculum-based versions of QMIX and VDN as a solution to these problems and compare their results with the standard algorithms. The results show a clear performance gain in terms of a common cumulative reward metric.

Read the paper · More papers on PaperTik