Predictive state representations with options
Britton Wolfe, Satinder Pal Singh · 2006
Recent work on predictive state representation (PSR) models has focused on using predictions of the outcomes of open-loop action sequences as state. These predictions answer questions of the form “What is the probability of seeing observation sequence o1, o2,..., oN if the agent takes action sequence a1, a2,..., aN from some given history?” We would like to ask more expressive questions in our representation of state, such as “If I behave according to some policy until I terminate, what will be my last observation?” We extend the linear PSR framework to answer questions like these about options – temporally extended, closed-loop courses of action – bounding the size of the linear PSR needed to model questions about a certain class of options. We introduce a hierarchical PSR (HPSR) that can make predictions about both options and primitive action sequences and show empirical results from learning HPSRs in simple domains. Existing work with predictive state representations (PSRs) focuses on using predictions about open-loop action sequences as state. These predictions answer questions of the form “What is the probability of seeing observation sequence o1, o2,..., oN if the agent takes action sequence a1, a2,..., aN in some given history?” Littman et al. (2002) showed that predictions of this form are sufficient in that they can perfectly capture state and can be used to make any prediction, i.e., answer any question, about the system. In general, the number of predictions in the state vector grows linearly with the number of underlying or hidden system states and this can be too large for practical purposes. Of course, if one truly wants a