On a Continuously Discounted Vector Valued Markov Decision Process
Kensuke Tanaka, Chikao Matsuda · Journal of Information and Optimization Sciences · 1990
The paper deals with continuous time vector valued Markov decision process on a general state space. The vector rewards are continuously discounted at rate α>0. The optimization criterion of the decision process is made from the domination structure determined by a given convex cone L’. By using an affine transformation of the reward in R m to R, we show the existence of an L’-optimal solution under some conditions and, then the relations between an L’-optimal solution and an optimal solution of real valued Markov decision process are characterized. Further, the supremum is considered in the case of which an optimal policy does not exist.