Scaling Up Approximate Value Iteration with Options: Better Policies with Fewer Iterations

Timothy Mann, Shie Mannor · 2014

We show how options, a class of control struc-tures encompassing primitive and temporally ex-tended actions, can play a valuable role in plan-ning in MDPs with continuous state-spaces. An-alyzing the convergence rate of Approximate Value Iteration with options reveals that for pes-simistic initial value function estimates, options can speed up convergence compared to plan-ning with only primitive actions even when the temporally extended actions are suboptimal and sparsely scattered throughout the state-space. Our experimental results in an optimal replace-ment task and a complex inventory manage-ment task demonstrate the potential for options to speed up convergence in practice. We show that options induce faster convergence to the op-timal value function, which implies deriving bet-ter policies with fewer iterations. 1.

Read the paper · More papers on PaperTik