Nonparametric Bayesian Policy Priors for Reinforcement Learning

Finale Doshi‐Velez, David Wingate, Nicholas Roy, Joshua B. Tenenbaum · DSpace@MIT (Massachusetts Institute of Technology) · 2010

We consider reinforcement learning in partially observable domains where the agent can query an expert for demonstrations. Our nonparametric Bayesian ap-proach combines model knowledge, inferred from expert information and inde-pendent exploration, with policy knowledge inferred from expert trajectories. We introduce priors that bias the agent towards models with both simple representa-tions and simple policies, resulting in improved policy and model learning. 1

Read the paper · More papers on PaperTik