STRUCTURE OF OPTIMAL POLICIES FOR DISCOUNTED SEMI-MARKOV DECISION PROGRAMING WITH UNBOUNDED REWARDS

Iu K · 1986

In this paper, we discuss the structure of optimal policies for discounted semi-Markov decision programing with unbounded rewards (abbreviated to URSMDP) discussed by Lippman. We prove that if a policy π~* = (π_0~*, π_1~*, π_2~*,…) is α-optimal, then the stochastic stationary policy π_0~(*∞)= (π_0~*, π_0~*,…) is a-optimal too (for the same α). For any given integer n≥1, we give the sufficient conditions in which π_n~*=(π_n~*,π_n~*,π_n~*,…) (under suitable history) is α-optimal. Any stochastic stationary α-optimal policy π_0~∞ can be decomposed into some optimal determinate stationary policies (maybe infinite), and it must be a convex combination of these optimal determinate stationary policies (for the same α).

Read the paper · More papers on PaperTik