A Least Squares Q-Learning Algorithm for Optimal Stopping Problems

Huizhen Yu, Dimitri P. Bertsekas · 2007

We consider the solution of discounted optimal stopping problems using linear function approximation methods. A Q-learning algorithm for such problems, proposed by Tsitsiklis and Van Roy, is based on the method of temporal dierences and stochastic approximation. We propose alternative algorithms, which are based on projected value iteration ideas and least squares. We prove the convergence of some of these algorithms and discuss their properties.

Read the paper · More papers on PaperTik