Multi-Agent Actor-Critic Multitask Reinforcement Learning based on GTD(1) with Consensus

Miloš S. Stanković, Marko Beko, Nemanja Ilić, Srdjan S. Stanković · 2022 IEEE 61st Conference on Decision and Control (CDC) · 2022

In this paper, a new distributed multi-agent off-policy Actor-Critic algorithm for collaborative multitask reinforcement learning is proposed. The Critic stage is based on the distributed gradient temporal difference algorithm GTD(1), while the Actor stage is derived from a predefined global criterion function and consists of a complementary consensus-based exact policy gradient algorithm. A proof that the Feller-Markov properties hold for the derived algorithm at the Actor stage is derived. The weak convergence of the algorithm to the set of stationary points of an attached ODE is proved under mild conditions using the two-time-scale stochastic approximation arguments. An experimental verification of the algorithm properties is given, demonstrating its high efficiency and practical applicability.

Read the paper · More papers on PaperTik