Decentralized Multi-Agent Multi-Task Q-Learning with Function Approximation for POMDPs

Miloš S. Stanković, Marko Beko, Srdjan S. Stanković · 2024

In this paper we propose a novel distributed gradient-based two-time-scale algorithm for decentralized multi-agent multi-task learning (MTL) using a linear approximation of the optimal action value function (Q -function) in POMDPs. The algorithm is based on the idea of using in a concurrent way recursive Bayesian state belief filters for estimation of the system model parameters, prediction of the hidden state and definition of the optimal approximation parameters of the local Q-functions. The main MTL algorithm is composed of: 1) local parameter updates based on an off-policy gradient-based learning algorithm with target policy belonging to the greedy or Gibbs classes, and 2) a linear stochastic time-varying consensus scheme for parameters shared between the agents in order to achieve the MTL goal. It is proved, under general assumptions, that the parameter estimates generated by the proposed algorithm weakly converge to a bounded invariant set of the corresponding ordinary differential equations (ODE). Simulation results illustrate the effectiveness of the algorithm.

Read the paper · More papers on PaperTik