A Note on the Convergence of Policy Iteration in Markov Decision Processes with Compact Action Spaces

A. Y. Golubin · Mathematics of Operations Research · 2003

The undiscounted, unichain, finite state Markov decision process with compact action space is studied. We provide a counterexample for a result in Hordijk and Puterman (1987) and give an alternate proof of the convergence of policy iteration under the condition that there exists a state that is recurrent under every stationary policy. The analysis essentially uses a two-term matrix representation for the relative value vectors generated by policy iteration procedure.

Read the paper · More papers on PaperTik