Guided Policy Search
Sergey Levine, Vladlen Koltun · 2013
Direct policy search can effectively scale to high-dimensional systems, but complex policies with hundreds of parameters often present a challenge for such methods, requir-ing numerous samples and often falling into poor local optima. We present a guided pol-icy search algorithm that uses trajectory op-timization to direct policy learning and avoid poor local optima. We show how differential dynamic programming can be used to gener-ate suitable guiding samples, and describe a regularized importance sampled policy opti-mization that incorporates these samples into the policy search. We evaluate the method by learning neural network controllers for planar swimming, hopping, and walking, as well as simulated 3D humanoid running. 1.