Generalized Classication-bas ed Approximate Policy Iteration
Amir‐massoud Farahmand, Doina Precup, Mohammad Ghavamzadeh, Marc Peter Deisenroth, Jan Peters · 2012
Classication-bas ed approximate policy iteration is very useful when the optimal policy is easier to represent and learn than the optimal value function. We theoretically analyze a general algorithm of this type. The analysis extends existing work by allowing the policy evaluation to be performed by any reinforcement learning algorithm, by handling nonparametric representations of policies, and by providing tighter convergence bounds. A small illustration shows that this approach can be faster than purely value-based methods.