Enhanced Proximal Policy Optimization for Complex Game AI: Applying Reinforcement Learning to Super Mario
<p>Lei Wang<sup>1</sup>, Bo Li<sup>2</sup>, Shengyu Wang<sup>3</sup>, Tingting Wang<sup>4</sup></p> · Academic Journal of Computing & Information Science · 2024
This paper presents an optimized implementation of Proximal Policy Optimization (PPO) for controlling an AI agent in the Super Mario environment. By introducing enhancements such as adaptive clipping, dual-clip objectives, and experience replay, our model addresses common limitations in standard PPO, such as unstable updates and sample inefficiency. Experimental results demonstrate that the enhanced PPO model achieves a completion rate exceeding 95% across Super Mario levels, utilizing fewer samples and exhibiting more stable convergence than baseline models. This study highlights the effectiveness of PPO in dynamic decision-making scenarios and provides a foundation for future reinforcement learning advancements.