Double Q($\\sigma$) and Q($\\sigma, \\lambda$): Unifying Reinforcement Learning Control Algorithms
Markus Dumke · arXiv (Cornell University) · 2017
Temporal-difference (TD) learning is an important field in reinforcement learning. Sarsa and Q-Learning are among the most used TD algorithms. The Q($\\sigma$) algorithm (Sutton and Barto (2017)) unifies both. This paper extends the Q($\\sigma$) algorithm to an online multi-step algorithm Q($\\sigma, \\lambda$) using eligibility traces and introduces Double Q($\\sigma$) as the extension of Q($\\sigma$) to double learning. Experiments suggest that the new Q($\\sigma, \\lambda$) algorithm can outperform the classical TD control methods Sarsa($\\lambda$), Q($\\lambda$) and Q($\\sigma$).