Bayesian Policy Gradient Algorithms
Mohammad Ghavamzadeh, Yaakov Engel · The MIT Press eBooks · 2007
Policy gradient methodsare reinforcement learningalgorithmsthat adapt a parameterizedpolicy by fo lowinga performance gradient estimate. Conventional policy gradient methodsuse Monte-Carlo techniquesto estimate thisgradient, and improve the policy by adjustingthe parametersin the direction of that gradient estimate. Since Monte Carlo methodstendto have highvariance, a large number of samplesisrequiredto attain accurate estimates, resultingin slow convergence. In thispaper, we propose a Bayesian frameworkthat modelsthe policy gradient asa Gausian proce s. Thisreducesthe number of samplesneededto obtain accurate gradient estimates. Moreover, estimatesof the naturalgradient anda measure of the uncertainty in the gradient estimatesare providedat little extra cost. 1