Total Expected Discounted Reward MDPS : Existence of Optimal Policies
Eugene Aleksandrovich Feinberg · Wiley Encyclopedia of Operations Research and Management Science · 2011
Abstract This article describes the results on the existence of optimal and nearly optimal policies for Markov Decision Processes (MDPs) with total expected discounted rewards. The problem of optimization of total expected discounted rewards for MDPs is also known under the name of discounted dynamic programming.