Bayes-Adaptive POMDPs: A New Perspective on the Explore-Exploit Tradeoff in Partially Observable Domains.
Joëlle Pineau, Stéphane Ross, Brahim Chaib-draa · ISAIM · 2008
Bayesian Reinforcement Learning has generated substantial interest recently, as it provides an elegant solution to the exploration-exploitation trade-off in reinforcement learning. However most investigations of Bayesian reinforcement learning to date focus on the standard Markov Decision Processes (MDPs). Our goal is to extend these ideas to the more general Partially Observable MDP (POMDP) framework, where the state is a hidden variable. This difficult decision-making problem can be formulated cleanly by simply extending the state to include the model parameters themselves. However closed-form solutions are not possible. This paper explores a family of approximations for solving this problem. These approaches are able to trade-off between (1) improving knowledge of the POMDP domain through interaction with the environment, (2) resolving uncertainty about the current state, and (3) choosing actions with high expected reward.