Hierarchical Sampling for Least-Squares Policy Iteration
Devin Schwab · OhioLink ETD Center (Ohio Library and Information Network) · 2016
For large Sequential Decision Making tasks, an agent may need to make lots of exploratory interactions within the environment in order to learn the optimal policy.Large amounts of exploration can be costly in terms of computation, time for interactions, and physical resources.This thesis studies approaches to incorporate prior knowledge to reduce the amount of exploration.Specifically, I propose an approach that uses a hierarchical decomposition of the Markov Decision Process to guide an agent's sampling process, in which the hierarchy is treated as a set of constraints on the sampling process.I show theoretically that, in terms of distributions of stateaction pairs sampled with respect to hierarchical states, variants of my approach have good convergence properties.Next, I perform an extensive empirical validation of my approach by comparing my methods to baselines which do not use the prior information during the sampling process.I show that using my approach, not only will irrelevant state-action pairs be avoided while sampling, but that the agent can learn a hierarchically optimal policy with far fewer samples than the baseline techniques.viii