Online Multi-Task Learning for Policy Gradient Methods
Haitham Bou Ammar, Eric R. Eaton, Paul Ruvolo, Matthew Edmund Taylor · 2014
Policy gradient algorithms have shown consider-able recent success in solving high-dimensional sequential decision making tasks, particularly in robotics. However, these methods often require extensive experience in a domain to achieve high performance. To make agents more sample-efficient, we developed a multi-task policy gra-dient method to learn decision making tasks con-secutively, transferring knowledge between tasks to accelerate learning. Our approach provides ro-bust theoretical guarantees, and we show empir-ically that it dramatically accelerates learning on a variety of dynamical systems, including an ap-plication to quadrotor control. 1.