Autonomous learning of POMDP state representations from surprises

Thomas Collins, Wei‐Min Shen · 2018

There is an ever-increasing need for autonomous robots that are capable of operating in a range of challenging environments exhibiting both partial-observability and stochasticity. Standard techniques for learning in such environments often require human-engineered features, most commonly a human-designed state space. Engineering these features demands extensive domain knowledge, and changes to the task or the agent often necessitate re-engineering. These limitations have given rise to end-to-end, predictive approaches, such as Predictive State Representations (PSRs) and our Stochastic Distinguishing Experiments (SDEs), that encode a representation of the agent's state in the probabilities of key sequences of actions and observations (i.e., experiments the agent can perform). The problem of discovering appropriate experiments has remained extremely challenging, in part because existing techniques treat them as decoupled from any latent structure in the agent's environment. In this paper, we extend our SDE representation into a hybrid latent-predictive representation of state that can provably model a useful subclass of POMDP environments exactly (and any POMDP environment approximately). We provide an active, incremental algorithm for autonomously learning such representations in unknown environments from experience. The key idea is that the agent begins using only its observations as a state space and splits those states into a hierarchy of additional latent states when it is surprised by the entropy resulting from the repeated executions of experiments that are automatically designed and selected based on these surprises to statistically disambiguate identical-looking states. The results of these experiments form unique predictive labels for each latent state. We present experimental results demonstrating the feasibility of this learning procedure. The corresponding expanded version of this paper provides the theoretical proofs of the representational capacity of this model.

Read the paper · More papers on PaperTik