Policy‐Gradient Algorithms
Olivier Buffet · 2013
This chapter focuses on gradient algorithms, called policy gradients, which have been particularly studied. After a brief reminder about the notion of gradient, it presents two types of approaches: direct policy-gradient algorithms and actor-critic policy gradients. The chapter presents examples of forms that the parameterized policy can take. It shows how the gradient can be estimated in different situations. The chapter shows how to calculate an estimate of the gradient of a differentiable function for a Markov Decision Process (MDP) with a time horizon of length l. It extends this computation to the case of an MDP of finite time horizon T. A motivation for employing direct policy search methods is to reduce the need for memory and computation time by limiting the search space to a sub-space of possible policies. Controlled Vocabulary Terms gradient methods; infinite horizon