Process-oriented planning and average-reward optimality

Craig Boutiller, Martin L. Puterman · 1995

hnp //www cs ubc ca/spider/cebty/craig html We argue that many AI planning problems should be viewed as process-onented, where the aim is to produce a policy or behavior strategy with no termination condition in mind, as opposed to goal-onented The full power of Markov decision models, adopted recently for AI planning, becomes apparent with process-onented problems The question of appropriate opdmallry criteria becomes more cnncal in this case, we argue that average reward optimahty is most suitable While construction of averageoptimal policies involves a number of subtleties and computational difficulties, certain aspects of the problem can be solved using compact action representations such as Bayes nets In particular, we provide an algorithm that identifies the structure of the Markov process underlying a planning problem- a crucial element of constructing average optimal policies- without explicit enumeration of the problem state space 1

Read the paper · More papers on PaperTik