Semi-Infinite Weighted Markov Decision Processes

Mohammed Abbad, Khalid Rahhali · Stochastic Models · 2008

In this article, weighted reward Markov decision processes (semi-infinite WMDP) with finite state and countable action spaces are considered. The “weighted reward” refers to an appropriately normalized convex combination of the discounted and long-run average reward criteria. This criterion allows the controller to trade-off short-term rewards versus long-run rewards. We prove that under some conditions, the supremum in the class of general strategies is equivalent to the supremum in the class of relatively simple “ultimately deterministic” strategies. These are strategies that behave just like deterministic stationary strategies, after a certain point in time. We present an iterative algorithm for computing a δ-optimal simple ultimately deterministic strategy. The steps of the algorithm are based on the one developed in Ref. for finite WMDP.

Read the paper · More papers on PaperTik