Federated TD Learning in Heterogeneous Environments with Average Rewards: A Two-timescale Approach with Polyak-Ruppert Averaging

Ankur Naskar, Gugan Thoppe, Abbasali Koochakzadeh, Vijay Gupta · 2024

Federated Reinforcement Learning (FRL) provides a promising way to speedup training in reinforcement learning using multiple edge devices that can operate in parallel. Recently, it has been shown that even when these edge devices have access to different dynamic models, an optimal convergence rate that has a linear speedup proportional to the number of devices is achievable. However, this result requires that the stepsize in the algorithm be chosen in a manner dependent on the unknown model parameters. Also, it applies only to a discounted setting, which has been argued to fit episodic tasks better than continuing control tasks. In this paper, we obtain finite-time bounds for heterogeneous FRL with average rewards. We show that the optimal convergence rate with a linear speedup is possible even with a universal stepsize choice, independent of the underlying dynamics. To achieve our result, we modify the existing one-timescale FRL method to a novel two-timescale variant that additionally incorporates iterate averaging.

Read the paper · More papers on PaperTik