Time-regularized interrupting options

Daniel J. Mankowitz, Timothy Mann, Shie Mannor · 2014

High-level skills relieve planning algorithms from low-level details. But when the skills are poorly designed for the domain, the resulting plan may be severely suboptimal. Sutton et al. (1999) made an important step towards resolv-ing this problem by introducing a rule that auto-matically improves a set of skills called options. This rule terminates an option early whenever switching to another option gives a higher value than continuing with the current option. How-ever, they only analyzed the case where the im-provement rule is applied once. We show condi-tions where this rule converges to the optimal set of options. A new interrupting Bellman operator that simultaneously improves the set of options is at the core of our analysis. One problem with the update rule is that it tends to favor lower-level skills. We introduce a regularization term that fa-vors longer duration skills. Experimental results demonstrate that this approach can derive a good set of high-level skills even when the original set of skills cannot solve the problem. 1.

Read the paper · More papers on PaperTik