Coaching: Human-assisted approach for reinforcement learning
Nakarin Suppakun, Thavida Maneewarn · 2017
The technique called ‘Coaching’ is proposed in this work. Coaching is a method to accelerate learning by employing a human knowledge at the early phase of learning. The human coach can guide a robot behavior by temporarily replacing the global goal with an intermediate target. During the coaching process, an action is chosen by a greedy policy such that it is most likely driving the robot to the intermediate target. When the intermediate target is reached, a normal pair of policy (f-greedy) and reward function is switched back. However, the global reward function is still used for updating the state-action value during both coaching and non-coaching periods. A human coach can guide the robot by using 8 verbal commands to place the intermediate target location relative to the agent current location. In this work, Q learning algorithm was used to test with the proposed method on 2 learning tasks: ball following, and obstacle avoidance. The proposed technique resulted in faster learning performance when compared to the traditional method of reinforcement learning.