Policy Improvement Directions for Reinforcement Learning in Reproducing Kernel Hilbert Spaces
Santiago Paternain, Juan Andrés Bazerque, Austin Small, Alejandro Ribeiro · 2019
The celebrated policy gradient algorithm solves the reinforcement learning problem by updating the control policy in the direction of the gradient of the value function. For the algorithm to converge it requires to be reset to the initial state after each policy update, or to perform exploring starts in which the system is reset to a random state. These restarts are possible in an offline setup where the policy is trained over multiple trajectories of the system. However, they prevent a fully online implementation in which the policy is optimized on the fly, that is, while the system is continuously following a single trajectory. In this work, we focus on the latter problem. We assume that the spaces of states and actions are continuous, and the policies are sought to be randomized versions of continuous functions belonging to a reproducing kernel Hilbert space. Under this framework, our main result is proving that gradients computed when the system is in a state down the trajectory serve as ascent directions for the value function defined with respect to the initial state. Building on that result we can prove convergence of the policy iterates to a ball of the critical points of the original value function. Numerical experiments in navigation problems support the theoretical conclusions.