Tuning Apex DQN: A Reinforcement Learning based Deep Q-Network Algorithm
Dhani Ruhela, Amit Ruhela · 2024
Technological inventions in the last decade have increasingly geared towards developing autonomous systems that belong to robotics, healthcare, games, smart grids, finance and other domains. Such advanced developments have become possible through the incredible capabilities of Deep Reinforcement Learning (DRL), which combines reinforcement learning (RL) and deep learning and resembles how humans begin learning from the world from the very moment they are born. In RL, a machine or an agent interacts with an environment by exploring or exploiting various policies to reach an outcome. The learning happens with each reward or punishment an agent receives, attaining a desirable or undesirable state. The agent aims to maximize the cumulative reward by interacting numerous times with the environment. Deep Q-Networks (DQN) is a powerful RL technique that aims to maximize the action-value function (also known as Q values) and achieve exceptional performance. Ray’s RLLib library, an open-source library, supports production-level, highly distributed RL algorithms while maintaining unified and straightforward APIs for various industry applications. However, it is incredibly challenging to optimize the performance of the RL algorithms, given the vast number of tuning knobs or parameters supported by the library. Understanding the behavior of these knobs demands a significant amount of computing resources that execute applications with a range of possible values. To the best of our belief, this is the first research work that exposes the critical parameters to tune the performance of the class of DQN algorithms in RLLib. A thorough investigation was performed with a famous Cartpole game that indicated over 2x performance improvements by selecting appropriate tuning knobs.