Recent Advances in Reinforcement Learning

Mathukumalli Vidyasagar · 2020

In this paper, we give a brief review of Markov Decision Processes (MDPs), and how Reinforcement Learning (RL) can be viewed as MDP where the parameters are unknown. Specific topics discussed include the Bellman equation and the Bellman operator, and value and policy iterations for MDPs, together with recent "empirical" approaches to solving the Bellman equation and applying the Bellman iteration. In addition to the well-established method of Q-learning, we also discuss the more recent approach known as Zap Q-learning.

Read the paper · More papers on PaperTik