Selective Representation of Multi-Agent Q Learning

Pengcheng Liu, Kangxin He, Hui Li · 2023

Value decomposition is a common approach to implementing centralized training and distributed execution paradigms in collaborative multi-agent reinforcement learning. By ensuring the IGM (Individual-Global-Max) principle (requiring consistency between joint and local action selection), value decomposition-based methods always perform excellently in many complex tasks. However, these methods restrict the Q-value of the joint action, making it impossible for the joint Q-value to fully express the true value function and thus achieve optimal consistency. Existing methods address the above problem by achieving a complete representation of the joint action-values, but these methods perform poorly in some complex tasks due to learning additional information or parameters. In this paper, we first consider the idea of a selective representation of value decomposition methods and propose the action-global-maximum (AGM) concept. Then, we propose a new value decomposition framework to reach optimal consistency by achieving a selective representation of the AGM. Our method focuses on avoiding the influence of joint Q-values outside the AGM on the value decomposition and guarantees the correct representation of the AGM. Extensive experiments in three different environments show that our approach achieves not only optimal consistency between individual and joint actions but also has good performance in most environments.

Read the paper · More papers on PaperTik