Yet More Optimistic Temporal Difference Learning for Game Mini2048
Kiminori Matsuzaki, Shunsuke Terauchi · 2024
abstract- Reinforcement learning is now an important method for developing strong computer players for games. One of the most fundamental issues in reinforcement learning is the exploration-exploitation dilemma. For Game 2048, existing reinforcement learning algorithms mostly pursued exploitation only, and learning with optimistic initialization has recently achieved state-of-the-art results. In this study, we investigate how and how much optimistic initialization contributes to exploration by using the results of the perfect analysis of Mini2048, a reduced variant of 2048. We find room for improvement in terms of exploration after the detailed analysis, and then we design two learning algorithms with exploratory move selection enhanced with two strategies. The results of the proposed training methods outperform learning with optimistic initialization only, especially when combined with deep search. We also discuss the applicability of our results to the original 2048.