Bayesian Policy Gradient Algorithms

Mohammad Ghavamzadeh, Yaakov Engel · The MIT Press eBooks · 2007

Policy gradient methodsare reinforcement learningalgorithmsthat adapt a parameterizedpolicy by fo lowinga performance gradient estimate. Conventional policy gradient methodsuse Monte-Carlo techniquesto estimate thisgradient, and improve the policy by adjustingthe parametersin the direction of that gradient estimate. Since Monte Carlo methodstendto have highvariance, a large number of samplesisrequiredto attain accurate estimates, resultingin slow convergence. In thispaper, we propose a Bayesian frameworkthat modelsthe policy gradient asa Gausian proce s. Thisreducesthe number of samplesneededto obtain accurate gradient estimates. Moreover, estimatesof the naturalgradient anda measure of the uncertainty in the gradient estimatesare providedat little extra cost. 1

Read the paper · More papers on PaperTik