QVA-learning for playing the game of Snake
Ilse Pubben · 2020
In this thesis we will introduce a new reinforcement learning algorithm, QVA-learning, which is a combination of QV-learning and advantage updating. We will test this algorithm on the game of Snake and compare its performance with Q-learning and QV-learning. The state will be represented with vision grids of size 3×3, 5×5 and 7×7. We will also make use of an MLP as function approximator. We found that overall QVA-learning did not perform better than Q-learning or QV-learning. We also found that QVA-learning started learning earlier than the other algorithms when using a vision grid of size 7×7 (i.e., with more input nodes).