Continuous Time Control of Markov Processes on an Arbitrary State Space: Discounted Rewards

Bharat T. Doshi · The Annals of Statistics · 1976

The paper deals with continuous time Markov decision processes on a fairly general state space. The rewards are continuously discounted at rate $\alpha > 0$. A set of conditions is shown to be necessary and sufficient for a policy to be optimal. For the special case of time independent reward function and under the assumption that the action space is finite a policy improvement algorithm is proposed and its convergence to an optimal policy is proved.

Read the paper · More papers on PaperTik