Two-Timescale Networks for Nonlinear Value Function Approximation
Wesley Chung · ERA: Education and Research Archive (University of Alberta) · 2019
Policy evaluation, learning value functions, is an integral part of the reinforcement learning problem. In this thesis, I propose a neural network architecture, the Two-Timescale Network (TTN), for value function approximation which utilizes linear function approximation for the value function with learned features. By separating these two learning processes—approximating the value function and learning features—we can utilize classic policy evaluation methods suited for linear function approximation but still obtain nonlinear estimates of the value function. Additionally, the separation facilitates proving convergence guarantees for the value estimates. This thesis contains empirical investigations about the choice of linear policy evaluation algorithm, the choice of objective for feature-learning and also presents some experiments in the control setting. We find that TTNs perform competitively with other algorithms which train both the features and the value function estimates jointly. In particular, utilizing least-squares temporal difference methods seem to provide the largest benefit and eligibility traces can also be helpful for linear time TD algorithms. Overall, this thesis provides evidence that separating feature and value learning is a promising direction for nonlinear value function approximation.