Learning Complex Neural Network Policies with Trajectory Optimization

Sergey Levine, Vladlen Koltun · 2014

Direct policy search methods offer the promise of automatically learning controllers for com-plex, high-dimensional tasks. However, prior ap-plications of policy search often required spe-cialized, low-dimensional policy classes, limit-ing their generality. In this work, we introduce a policy search algorithm that can directly learn high-dimensional, general-purpose policies, rep-resented by neural networks. We formulate the policy search problem as an optimization over trajectory distributions, alternating between opti-mizing the policy to match the trajectories, and optimizing the trajectories to match the policy and minimize expected cost. Our method can learn policies for complex tasks such as bipedal push recovery and walking on uneven terrain, while outperforming prior methods. 1.

Read the paper · More papers on PaperTik