Hierarchical Optimal Control of MDPs
Amy McGovern, Doina Precup, Balaraman Ravindran, Satinder Singh, Stanley M. Sutton, Scott F. Richard · 1998
Fundamental to reinforcement learning, as well as to the theory of systems and control, is the problem of represent-ing knowledge about the environment and about possible courses of action hierarchically, at a multiplicity of interre-lated temporal scales. For example, a human traveler must decide which cities to go to, whether to fly, drive, or walk, and the individual muscle contractions involved in each step. In this paper we survey a new approach to reinforce-ment learning in which each of these decisions is treated uniformly. Each low-level action and high-level course of action is represented as an option, a (sub)controller and a termination condition. The theory of options is based on the theories of Markov and semi-Markov decision pro-cesses, but extends these in significant ways. Options can be used in place of actions in all the planning and learn-ing methods conventionally used in reinforcement learning. Options and models of options can be learned for a wide variety of different subtasks, and then rapidly combined to solve new tasks. Options enable planning and learning si-multaneously at a wide variety of times scales, and toward a wide variety of subtasks, substantially increasing the ef-ficiency and abilities of reinforcement learning systems.