Performance Reliability of Reinforcement Learning Algorithms in Obstacle Avoidance Game with Differing Reward Formulations

Brandon Hansen, Haley R. Dozier · 2022

When formulating environments for complex application areas, using analogies to games is beneficial as they provide convenient models to test algorithm performance in ways that are transferable to realistic environments. We propose a Frogger like grid based environment containing a simple action space, dynamic obstacles, and discrete game loop for testing Proximal Policy Optimization 2 and Deep Q-Network with comparisons to a heuristic and random agent. The environment contains four different reward function implementations, along with two different environment variations to explore adaptability. Seeing how these different parameters effect not just the average performance of the algorithm, but also the reliability of the performance is of concern as reliability determines the expectations of any single performance of an reinforcement learning agent. Experiments in these environments demonstrate common behaviors of reinforcement learning algorithms showing possible strengths and weakness of the approaches when applied to more complex decision-making scenarios. We explore these behaviors through evaluation techniques meant to measure the agents accumulation of reward in the game. Cross comparing these evaluation techniques elucidates the causation behind agent performance.

Read the paper · More papers on PaperTik