Factored Markov Decision Processes

Thomas Degris, Olivier Sigaud · 2013

This chapter describes Factored Markov Decision Processes (FMDPs) first proposed by [BOU 95, BOU 99]. The chapter describes the framework and how problems are modeled. It describes different planning methods able to take advantage of the structure of the problem to calculate optimal or near-optimal solutions. To illustrate the FMDP framework, the chapter uses a well-known example in the literature named Coffee Robot. Using this example, it describes the decomposition of the transition and the reward functions with a formalization of function-specific independencies. The chapter proposes a formalization of context-specific independencies. In FMDPs, function-specific independencies are formalized with dynamic Bayesian networks. In addition to using function-specific independence, Structured Value Iteration (SVI) and Structured Policy Iteration (SPI) exploit context-specific independence by using decision trees to represent the different functions of the problem. Controlled Vocabulary Terms belief networks; iterative methods; robots

Read the paper · More papers on PaperTik