From non-deterministic to probabilistic planning with the help of statistical relational learning

Ingo Thon, Bernd Gutmann, Martijn van Otterlo, Niels Landwehr, Luc De Raedt · Lirias (KU Leuven) · 2009

Using machine learning techniques for planning is getting in-creasingly more important in recent years. Various aspects of action models can be induced from data and then exploited for planning. For probabilistic planning, natural candidates are learning of action effects and their probabilities. For ex-pressive formalisms such as PPDDL, this is a difficult prob-lem since they can introduce easily a hidden data problem; the fact that multiple action outcomes may have generated the same experienced state transitions in the data. Further-more the action effects might be factored such that this prob-lem requires solving a constraint satisfaction problem within an expectation maximization scheme. In this paper we out-line how to utilize recent techniques from the field of statis-tical relational learning for this problem. More specifically, we show how techniques developed for the CPT-L model of relational probabilistic sequences can be applied to the prob-lem of learning probabilities in a PPDDL model. A CPT-L model concisely specify a Markov chain over arbitrary num-bers of objects in the domain and simultaneous applications of multiple actions. The use of efficient BDD-style represen-tations allows for fast and efficient learning in such domains. Even efficient online learning is possible as we will show in this paper. We relate to other learning approaches for similar domains and highlight the opportunities for incorporating our approach into architectures that can plan, execute the plan, and learn from the outcomes, in an online and incremental fashion.

Read the paper · More papers on PaperTik