Reinforcement Learning
Olivier Sigaud, Frédérick Garçia · 2013
This chapter presents reinforcement learning methods, where the transition and reward functions are not known in advance. Before presenting the main concepts of reinforcement learning, it gives a brief overview of the successive stages of research that led to the current formal understanding of the domain from the computer science viewpoint. Reinforcement learning stands at the intersection between the field of dynamic programming and the field of machine learning. In order to determine the best possible action in any situation, the agent must try a lot of actions, through a trial-and-error process that implies some exploration, giving rise to the exploration/exploitation trade-off that is central to reinforcement learning. Dynamic programming algorithms apply when the agent knows the transition and reward functions. By contrast, Monte Carlo methods do not require any knowledge of the transition and reward functions, but they are not local. Controlled Vocabulary Terms dynamic programming; learning (artificial intelligence); Monte Carlo methods