Human-assisted reinforcement learning demonstrated on the Flappy Bird Game

Jana Ristovska, Domen Šoberl · 2023

In model-free reinforcement learning, the agent usually starts learning by blind exploration, which can take a significant amount of time before starting to experience positive reinforcement that drives further progress.In this paper, we address the following question: Can the learning time be reduced if the agent first observes successful behaviors demonstrated by humans, before commencing its independent learning?We propose an adaptation of the traditional Q-learning algorithm, so that it can gradually integrate recorded demonstrations into the learning process, and demonstrate this method on the well-known Flappy Bird game.We recorded 1496 gameplays of Flappy Bird played by 22 volunteers, and selected 941 successful recordings to assist the Q-learning algorithm.The results show that such human assistance speeds up the learning process in a logarithmic manner, which means that the biggest gain is made in the initial stages of learning and becomes almost negligible later on.We show experimentally that in the initial stages of learning, human demonstrations contain more useful information than the agent can acquire independently.

Read the paper · More papers on PaperTik