Learning of non-parametric control policies with high-dimensional state features
Herke van Hoof, Jan Peters, Gerhard M. Neumann · Lincoln Repository (University of Lincoln) · 2015
Learning complex control policies from highdimensional sensory input is a challenge forreinforcement learning algorithms. Kernel methods that approximate values functionsor transition models can address this problem. Yet, many current approaches rely oninstable greedy maximization. In this paper, we develop a policy search algorithm thatintegrates robust policy updates and kernel embeddings. Our method can learn nonparametriccontrol policies for infinite horizon continuous MDPs with high-dimensionalsensory representations. We show that our method outperforms related approaches, andthat our algorithm can learn an underpowered swing-up task task directly from highdimensionalimage data.