Maintaining predictions over time without a model
Erik Talvitie, Satinder Pal Singh · 2009
A common approach to the control problem in par-tially observable environments is to perform a di-rect search in policy space, as defined over some set of features of history. In this paper we con-sider predictive features, whose values are condi-tional probabilities of future events, given history. Since predictive features provide direct information about the agent’s future, they have a number of ad-vantages for control. However, unlike more typical features defined directly over past observations, it is not clear how to maintain the values of predictive features over time. A model could be used, since a model can make any prediction about the future, but in many cases learning a model is infeasible. In this paper we demonstrate that in some cases it is possible to learn to maintain the values of a set of predictive features even when a learning a model is infeasible, and that natural predictive features can be useful for policy-search methods. 1