Action elimination and stopping conditions for reinforcement learning
Eyal Even-Dar, Shie Mannor, Yishay Mansour · 2003
We consider incorporating action elimination procedures in reinforcement learning algo-rithms. We suggest a framework that is based on learning an upper and a lower estimates of the value function or the Q-function and eliminating actions that are not optimal. We provide a model-based and a model-free vari-ants of the elimination method. We fur-ther derive stopping conditions that guar-antee that the learned policy is approxi-mately optimal with high probability. Sim-ulations demonstrate a considerable speedup and added robustness. 1.