Distributed Consensus-Based Multi-Agent Off-Policy Temporal-Difference Learning
Miloš S. Stanković, Marko Beko, Srdjan S. Stankovic · 2021 60th IEEE Conference on Decision and Control (CDC) · 2021
In this paper we propose two novel distributed consensus-based temporal-difference algorithms for multi-agent off-policy learning of linear approximation of the value function in Markov decision processes. The algorithms are composed of: 1) local parameter updates based on single-agent off-policy algorithms TD(λ) and ETD(λ), and 2) a linear dynamic consensus scheme. The algorithms are completely decentralized, allowing: 1) efficient parallelization and 2) applications in which all the agents may have completely different behavior policies and different initial state distributions while evaluating a single target policy. Starting from the properties of the underlying Feller-Markov processes, we show that, under nonrestrictive assumptions, the algorithms weakly converge to a unique consensus point. A discussion is given on the asymptotic parameter values at consensus, including estimation bias and variance. The algorithms’ properties are illustrated by characteristic simulation results.