Multi-Agent Recurrent Deterministic Policy Gradient with Inter-Agent Communication
Joohyun Cho, Mingxi Liu, Yi Zhou, Rong‐Rong Chen · 2023
In this paper, we introduce a novel approach to multi-agent coordination under partial state and observation, called Multi-Agent Recurrent Deterministic Policy Gradient with Differentiable Inter-Agent Communication (MARDPG-IAC). In such environments, it is difficult for agents to obtain information about the actions and observations of other agents, which can significantly impact their learning performance. To address this challenge, we propose a recurrent structure that accumulates partial observations to infer the hidden information and a communication mechanism that enables agents to exchange information to enhance their learning effectiveness. We employ an asynchronous update scheme to combine the MARDPG algorithm with the differentiable inter-agent communication algorithm, without requiring a replay buffer. Through a case study of building energy control in a power distribution network, we demonstrate that our proposed approach outperforms conventional Multi-Agent Deep Deterministic Policy Gradient (MADDPG) that relies on partial state only.