A Weighted Critic Update Approach to Multi Agent Twin Delayed Deep Deterministic Algorithm
Tamal Sarkar, Shobhanjana Kalita · 2021 IEEE 18th India Council International Conference (INDICON) · 2021
Multi-Agent Reinforcement Learning (MARL) has gained a lot more focus in the recent years because of its widespread use in multi-agent systems like autonomous vehicles, smart grids, traffic control management. Despite seeing significant developments in algorithmic design, problems such as overestimation of bias and instability in learning are still prevalent in MARL algorithms. Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm shapes policies of every agent irrespective of the setting- i.e. fully cooperative, fully competitive, and mixed setting. In its attempt to maximise the overall return of all agents, it suffers from bias overestimation which subsequently slows the convergence. Multi-Agent Twin Delayed Deep Deterministic (MATD3) was proposed to overcome the bias overestimation problem encountered in MADDPG algorithm. In this paper we propose a modification for the critic loss calculation of MATD3 algorithm which when applied, performs better than the original MATD3 algorithm. The proposed algorithm, termed MATD3-WCU in the paper, further stabilises the learning process and speeds up convergence. The proposed algorithm is tested on three scenarios of the state-of-art Multi-agent Particle Environment (MPE). The results against the experiment demonstrate that the MATD3-WCU algorithm performs better and provides a more stabilised learning with respect to the original MATD3 algorithm.