The Determination of Approximately Optimal Policies in Markov Decision Processes by the Use of Bounds
D. J. White · Journal of the Operational Research Society · 1982
In the general area of Markov decision processes, a lot of attention has been given to deriving upper and lower bounds for approximating the optimal performance level. These are, in themselves, not useful unless they can be used to derive an approximately optimal policy. The existing literature does this specifically in the context of the computational methods being used at the time. However, irrespective of the method used, a straight application of one step of Howard's policy space method will give the desired results.