Distributed Policy Evaluation with Local Updates over Time-Varying Communication Network

Matthew Crespo, Chinwendu Enyioha · 2025

This paper studies a consensus-based policy evaluation algorithm in a cooperative team of heterogeneous learners. To improve each agent’s approximation of their value function, they each update their weight parameters, then perform consensus updates with neighboring agents over a dynamic communication network. We present the analysis of an efficient fully distributed algorithm for cooperatively evaluating and optimizing their value functions over a time-varying communication network in the average-reward setting. Our main result shows that, using this algorithm, the agents can approximate the true value function, with the approximation error dependent on the number of communication rounds and number of time steps it takes for the communication network to be connected. To validate the theoretical results, we present accompanying numerical experiments that show the theoretical error bounds are not only tight, but get tighter as the time steps required to guarantee network connectivity reduces.

Read the paper · More papers on PaperTik