Agnostic KWIK learning and efficient approximate reinforcement learning

István Szita, Csaba Szepesvári · Conference on Learning Theory · 2011

A popular approach in reinforcement learning is to use a model-based algorithm, i.e., an algorithm that utilizes a model learner to learn an approximate model to the environment. It has been shown that such a model-based learner is ecient if the model learner is ecient in the so-called \knows what it knows (KWIK) framework. A major limitation of the standard KWIK framework is that, by its very denition, it covers only the case when the (model) learner can represent the actual environment with no errors. In this paper, we study the agnostic KWIK learning model, where we relax this assumption by allowing nonzero approximation errors. We show that with the new denition an ecient model learner still leads to an ecient reinforcement learning algorithm. At the same time, though, we nd that learning within the new framework can be substantially slower as compared to the standard framework, even in the case of simple learning problems.

Read the paper · More papers on PaperTik