(Machine) Learning Robot Control Policies
Daniel H. Grollman, Odest Chadwicke Jenkins · 2009
It currently requires years of education and practice before a skilled user can successfully program a sophisticated robot platform to perform a given task. We are exploring ways in which statistical machine learning techniques can enable Learning from Demonstration, an approach where users ‘reprogram ’ a robot without writing code. In this scenario, a user demonstrates the desired task and the robot learns to perform the task by observing its performance. We treat this learning as a form of Policy Transfer, where the decision making policy latent in the demonstrator is transitioned onto the robot. Taking perception and motion processes as fixed, we represent each policy as a functional mapping from perceived states to desired actions (π(ˆs) → a ∗). Using teleoperation, a demonstrator guides the robot through an instance of the desired behavior, creating a set of matched inputs and outputs. Function approximation techniques can then be applied to find an approximation of the control policy (ˆπ). We have left the tasks undefined, as we are interested in how robots can be made to learn Unknown Tasks, tasks not predefined during construction and original programming. Robots that exhibit Lifelong Learning, learning over extended periods (years) and in multiple domains, will likely need to deal with this issue. We have thus been exploring nonparametric function approximators. In addition, by using an algorithm capable of fast ( ∼ 30Hz) inference and prediction on our system, we can enable interactive tutelage, where