Primal-Dual Multi-Agent Trust Region Policy Optimization for Safe Multi-Agent Reinforcement Learning

Jie Li, Junjie Fu · 2024

In the field of multi-agent reinforcement learning (MARL), achieving high performance is crucial for the success of multi-robot systems. Meanwhile, avoiding unsafe behaviors is becoming a practical problem that needs to be addressed. However, ensuring safety in MARL remains challenging due to the necessity for each agent to not only ensure its own safety but also consider the safety of other agents to ensure overall safe team behavior. In this study, we propose a novel safe multi-agent reinforcement learning algorithm called MATRPO-Lagrangian, which addresses the multi-agent constrained policy optimization problem using a combination of the primal-dual method and trust region policy optimization. Experimental results on the safe MARL benchmark Safe Multi-Agent MuJoCo show that our method achieves competitive performance and significant safe constraint satisfaction ability compared to existing methods.

Read the paper · More papers on PaperTik