Distributed Multi-Agent Gradient Based Q-Learning with Linear Function Approximation
Miloš S. Stanković, Marko Beko, Srdjan S. Stanković · 2024
In this paper we propose a novel distributed gradient-based two-time-scale algorithm for multi-agent off-policy learning of linear approximation of the optimal action-value function (Q-function) in Markov decision processes (MDPs). The algorithm is composed of: 1) local parameter updates based on an off-policy gradient temporal difference learning algorithm with target policy belonging to either the greedy or the Gibbs distribution class and stationary behavior policies possibly different for each agent, and 2) a linear stochastic time-varying consensus scheme. It is proved, under general assumptions, that the parameter estimates generated by the proposed algorithm weakly converge to a bounded invariant set of the corresponding ordinary differential equation (ODE). Simulation results illustrate effectiveness of the proposed algorithm.