Learning mixed behaviours with parallel Q-learning
Guillaume J. Laurent, Emmanuel Piat · 2003
This paper presents a reinforcement learning algorithm based on a parallel approach of the Watkins's Q-learning. This algorithm is used to control a two axis micro-manipulator system. The aim is to learn complex behaviour such as reaching target positions and avoiding obstacles at the same time. The simulations and the tests with the real manipulator show that this algorithm is able to learn simultaneously opposite behaviours and that it generates interesting action policies with regard to global path optimization.