Comparison of reinforcement learning in game AI

Jintaro Nogae, Kanemitsu Ootsu, Takashi Yokota, Shun Kojima · 2022

In recent years, various reinforcement learning algorithms have been developed, and research is being conducted to apply them to game AI. We aim to develop a game AI that is strong against any game. To achieve this, we need to investigate existing reinforcement learning algorithms, to evaluate the maximum score, average score and standard deviations, and clarify their characteristics.This paper compares 19 typical games of Atari video games with on-policy algorithm PPO and off-policy algorithms DQN and ACER with average, maximum score, and standard deviation. As a result of the experiment, we find that it can be roughly divided into an algorithm (DQN) having maximum score and large standard deviation value and an algorithm (PPO) having low score and small standard deviation. In addition, ACER stops learning while the score is low for many games. We also find that DQN can achieve higher score than ACER and PPO in total. In addition, there are games with maximum score and large standard deviation value, and games with low maximum score but a low standard deviation. From the experimental results, we investigate how to realize general-purpose strong reinforcement learning game AI for the general public by using the rule of getting high score only once or Always get a high score.

Read the paper · More papers on PaperTik