Deep Q-Networks Applied to Blackjack
Piyush Rai · 2024
This study explores the application of Deep Q-Networks (DQN) to the game of Blackjack, aiming to understand the efficacy of reinforcement learning methods in a probabilistic environment with a natural reward scheme. Building on Mnih et al.’s seminal work, which introduced DQNs for large state and action spaces, I implement and train a DQN on a Blackjack state machine. My network evaluates five legal actions and is compared to a traditional Q-network trained over 50 million episodes. I analyzed my model’s performance using policy score and expected return metrics, finding that while my DQN approach significantly outperforms traditional Q-learning, it falls short of optimal policy performance. The results highlight the challenges in learning probabilistic transitions and rewards, suggesting that further tuning and advanced methods could improve policy accuracy. Despite these challenges, the study confirms that deep Q-networks represent a powerful tool for reinforcement learning applications.