NON-DISCOUNTED OPTIMAL POLICIES IN CONTROLLED MARKOV SET-CHAINS
Masanori Hosaka, Masami Kurano · Journal of the Operations Research Society of Japan · 1999
In a controlled Markov set-chain with a discount factor β, we consider the case of β = 1 as a limiting case of β < 1 and find a non-discounted optimal policy which maximizes Abel-sum of rewards in time over all stationary policies under some partial order. We analyze the behavior of discounted total rewards as discount factor β approaches 1 under regularity conditions, and prove the existence of a non-discounted optimal policy, applying the Kakutani's fixed point theorem and policy improvement method. As a numerical example the Toymaker's problem is considered.