Controlling Cardea: Fast Policy Search in a High Dimensional Space
Martin C. Martin · 2004
The essential dynamics algorithm is a novel policy search algorithm for learning in a class of stochastic Markov decision processes (MDPs) with continuous state and action spaces. We apply it to the control of a 5 degree of freedom robot arm atop a Segway base. Movement of the arm causes the base to translate and tilt, which in turn affects the movement of the arm. The state space has 14 dimensions, and the action space 5 dimensions, twice the dimensionality of typical policy search applications. Despite the highly non-linear dynamics, the algorithm is able to control the robot through a wide range. What’s more, this is accomplished using very little domain knowledge and no knowledge of dynamics. 1