On occupation measures for total-reward MDPs
Eric V. Denardo, Eugene Aleksandrovich Feinberg, Uriel G. Rothblum · 2008
This paper is based on our recent contribution that studies Markov decision processes (MDPs) with Borel state and action spaces and with the expected total rewards. The initial state distribution is fixed. According to, for a given randomized stationary policy, its occupation measure as a convex combination of occupation measures for simpler policies. If this is possible for a given policy, we say that the policy can be split. In particular, we are interested in splitting a randomized stationary policy into (nonrandomized) stationary policies or into a randomized stationary policies that are nonrandomized on a given subset of states. Though studies Borel-state MDPs with expected total rewards, some of its results are new for finite state and action discounted MDPs. This paper focuses on these results.