Policies that generalize: solving many planning problems with the same policy
Blai Bonet, Héctor Geffner · 2015
We establish conditions under which memory-less policies and finite-state controllers that solve one partially observable non-deterministic problem (PONDP) generalize to other problems; namely, problems that have a similar structure and share the same action and observation space. This is relevant to generalized planning where plans that work for many problems are sought, and to transfer learn-ing where knowledge gained in the solution of one problem is to be used on related problems. We use a logical setting where uncertainty is represented by sets of states and the goal is to be achieved with cer-tainty. While this gives us crisp notions of solution policies and generalization, the account also ap-plies to probabilistic PONDs, i.e., Goal POMDPs. 1