Error bounds for approximate policy iteration
Rémi Munos · 2003
In Dynamic Programming, convergence of algorithms such as Value Iteration or Policy Iteration results in discounted problems from a contraction property of the back-up operator, guaranteeing convergence to its xedpoint.