Markov Decisions on a Partitioned State Space

J. L. Smith · IEEE Transactions on Systems Man and Cybernetics · 1971

An important practical constraint on admissible control policies is defined for the Markov decision process. The framework of an algorithm based on the infinite return optimization algorithms of Howard and Jewell is suggested to compute the optimal policy under this constraint. Iterative convergence to the optimal policy cannot be guaranteed, but techniques proposed for state-space reduction and rapid resolution of undetermined policies should render many problems tractable.

Read the paper · More papers on PaperTik