Decentralized Deterministic Multi-Agent Reinforcement Learning
Antoine Grosnit, Desmond W. H. Cai, Laura Wynter · 2021 60th IEEE Conference on Decision and Control (CDC) · 2021
We provide a provably-convergent decentralized actor-critic algorithm for learning deterministic policies on continuous action spaces. Deterministic policies are important in real-world settings. To handle the lack of exploration inherent in deterministic policies, we consider both off-policy and on-policy settings. We give the expression of a local deterministic policy gradient, decentralized deterministic actor-critic algorithms and convergence guarantees for linearly-approximated value functions. This work will help enable decentralized MARL in high-dimensional action spaces and pave the way for more widespread use of MARL.