A Markov Decision Process (MDP)

Theodore J. Sheskin · 2014

A Markov decision process (MDP) is a sequential decision process for which the decisions produce a sequence of Markov chains with rewards or MCRs. As the introduction to Chapter 4 indicates, when decisions are added to a set of MCRs, the augmented system is called an MDP. An MDP generates a sequence of states and an associated sequence of rewards as it evolves over time from state to state, governed by both its transition probabilities and the series of decisions made. A rule that prescribes a set of decisions for all states is called a policy. When the planning horizon is infi nite, a policy is assumed to be independent of time, or stationary. When the planning horizon is infi - nite, the objective of an MDP model is to determine a stationary policy that is optimal in the sense that it will either maximize the gain, or expected reward per period, or maximize the expected total discounted reward received in every state. As in Chapter 4, both transition probabilities and rewards are assumed to be stationary.

Read the paper · More papers on PaperTik