Learning measures of progress for planning domains

Sung‐Wook Yoon, Alan Fern, Robert L. Givan · 2005

We study an approach to learning heuristics for planning do-mains from example solutions. There has been little work on learning heuristics for the types of domains used in determin-istic and stochastic planning competitions. Perhaps one rea-son for this is the challenge of providing a compact heuristic language that facilitates learning. Here we introduce a new representation for heuristics based on lists of set expressions described using taxonomic syntax. Next, we review the idea of a measure of progress (Parmar 2002), which is any heuris-tic that is guaranteed to be improvable at every state. We take finding a measure of progress as our learning goal, and describe a simple learning algorithm for this purpose. We evaluate our approach across a range of deterministic and stochastic planning-competition domains. The results show that often greedily following the learned heuristic is highly effective. We also show our heuristic can be combined with learned rule-based policies, producing still stronger results.

Read the paper · More papers on PaperTik