Finite-Sample Analysis of Multi-Agent Policy Evaluation with Kernelized Gradient Temporal Difference

Paulo Heredia, Shaoshuai Mou · 2020

In this work we will provide a finite-sample analysis of a distributed gradient temporal difference algorithm for policy evaluation with value functions that lie in Reproducing Kernel Hilbert Spaces (RKHS). This work focuses on multi-agent systems where each agent observes a private reward and agents can only communicate with nearby neighbors under time varying networks. The main result is a time-evolving upper bound of the second order error statistics of the algorithm, which accounts for the evolution of the consensus error as well as the average approximation error. This result shows that the distributed learning algorithm under consideration can achieve a bounded final error covariance that is inversely proportional to the algorithm step-size, which is consistent with results in the more general field of stochastic approximation.

Read the paper · More papers on PaperTik