Mastering Super Mario Bros.: Overcoming Implementation Challenges in Reinforcement Learning with Stable- Baselines3
Manav Khambhayata, Shivansh Singh, Neetu Bala, Arya Bhattacharyya · 2024
In the ever-evolving landscape of artificial intelligence, the application of reinforcement learning (RL) techniques to game playing has emerged as a captivating frontier, showcasing the capacity of intelligent agents to master complex environments and tasks autonomously. This research paper tackles the intricate process of implementing Reinforcement Learning (RL) algorithms for training agents in playing “Super Mario Bros.” within the OpenAI Gym environment, amidst challenges posed by version inconsistencies in libraries. We offer a focused approach, emphasizing the utilization of the latest versions of libraries such as OpenAI Gym and Stable-Baselines3 in PyTorch. By meticulously addressing errors stemming from version disparities, we provide a systematic guide to navigate through the implementation process successfully. Our methodology centers on the Proximal Policy Optimization (PPO) algorithm, known for its efficiency and reliability. Through iterative training, RL agents adeptly learn to maneuver the complexities of the game, optimizing strategies while surmounting obstacles. This paper presents key insights garnered throughout the implementation journey, including strategies for troubleshooting version inconsistencies and effectively mitigating errors. Serving as a practical resource for researchers and practitioners, our work aims to facilitate advancements in RL-based gaming AI research by offering a functional solution to address version inconsistencies and propel further development in the field.