On planning, prediction and knowledge transfer in fully and partially observable markov decision processes

Pablo Samuel Castro · 2011

uncertainty in large systems. The formalisms used to study this problem are fully and partially observable Markov Decision Processes (MDPs and POMDPs, respectively). The first contribution of this dissertation is a theoretical analysis of the behavior of POMDPs when only subsets of the observation set are used. One of these subsets is used to update the agent's state estimate, while the other subset contains observations the agent is interested in predicting and/or optimizing. The behaviors are formalized as three types of equivalence relations. The first groups states based on their values under optimal or general policies; the second groups states according to their ability to predict observations sequences; the third type is based on bisimulation, which is a well known equivalence relation borrowed from concurrency theory. Bisimulation relations can be generalized to bisimulation metrics. This dissertation introduces bisimulation metrics for an MDP with temporally extended actions (formalized as options) and proposes a new bisimulation metric that provides a tighter bound on the difference in optimal values. A new proof is provided for the convergence of an approximation method for computing bisimulation metrics that is based on statistical sampling, using only a finite number of samples. The new proof allows one to determine the minimum number of samples needed in order to achieve the desired quality of approximation with high probability. Although bisimulation metrics have been previously used for state space compression, this dissertation proposes using them to transfer policies from one MDP to another. In contrast to existing transfer work, the mapping between the two systems is determined automatically by means of the bisimulation metrics. Theoretical results are provided that bound the loss in optimality incurred by the transferred policy. A number of algorithms are introduced which are evaluated empirically in the context of planning and learning.

Read the paper · More papers on PaperTik