Comparing the Effectiveness of PPO and its Variants in Training AI to Play Game

Luobin Cui, Ying Gina Tang · 2023

Automated game intelligence is a crucial step in rapid game development. A promising research direction for automated game intelligence is reinforcement learning, and specifically, the proximal policy optimization (PPO) algorithm. Two variants of the PPO, Maskable PPO and Recurrent PPO, further extend the capabilities of the PPO. We compare the performance of these three algorithms in the 2D game Mario and a 3D car racing game environment. We also evaluate their performance and applicability by comparing the experimental results of the original algorithm authors. With our results, we provide recommendations on PPO configuration depending on the target game type, providing future developers with a benchmark to help them decide which algorithm is most applicable for their applications.

Read the paper · More papers on PaperTik