Value-Decomposition Multi-Agent Proximal Policy Optimization

Yanhao Ma, Jie Luo · 2022

We explore the policy-based approach for the newly well-liked centralized training and decentralized execution (CTDE) mechanism’s multi-agent reinforcement learning (MARL) job. Multi-agent proximal policy optimization (MAPPO) achieved optimal effect in multiple multi-agent cooperative tasks. However, it performs poorly in more complex multi-agent cooperative tasks because it does not solve the problem of credit assignment. To deal with this issue, we combine the method of value decomposition in value-based reinforcement learning with policy-based MAPPO, and propose a new actor-critic algorithm, namely value-decomposition multi-agent proximal policy optimization (VDPPO). VDPPO uses value decomposition to provide different rewards for agents to distinguish their contributions and train the policy network using the advantage function. To determine the effectiveness of our algorithm, we modify the multi-agent particle environments (MPE) and carry out experiments on it. The experimental findings demonstrate that our suggested method is more advantageous compared to several baseline algorithms and is suitable for more complex multi-agent cooperative tasks.

Read the paper · More papers on PaperTik