Reinforcement learning and mistake bounded algorithms
Yishay Mansour · 1999
Markov Decision Process (MDP) and Partially Observable MDP (POMDP) have become the model of choice in reinforcement learning.This work explores an interesting connection between mistake bounded learning algorithms and computing a near-best strategy, from a restricted class of strategies, for a given POMDP.We show that if a class of strategies has a mistake bound algorithm that makes at most d mistakes, then there is an algorithm to compute a near-best strategy from the class in time polynomial in l/c, the accuracy parameter, log(1/6), the confidence parameter, H, the horizon parameter, and exponential in d, the mistake bound.Our transformation assumes only the ability to execute actions in the POMDP and the ability to reset the POMDP to its initial state.