Self-play deep learning for games: Maximising experiences
Joseph West · Queensland University of Technology · 2020
This thesis describes several studies focused on improving the learning efficiency to train a combined tree-search/neural-network reinforcement learning agent for different board games. The work's primary contribution is a new approach to creating training experiences by enforcing a structured learning paradigm called the end-game-first curriculum which is shown to improve the speed of learning when compared against the current state-of-the-art agent. The thesis identifies a bottleneck in the self-play experience generation for a reinforcement learning agent and explores different methods to minimise the creation of poor experiences and maximise the use of experiences that are created.