Gradient-based relational reinforcement learning of temporally extended policies
Charles Gretton · 2007
1 We consider the problem of computing general policies for decision-theoretic planning problems with temporally extended rewards. We consider a gradient-based approach to relational reinforcement-learning (RRL) of policies for that setting. In particular, the learner optimises its behaviour by acting in a set of problems drawn from a target domain. Our approach is similar to inductive policy selection because the policies learnt are given in terms of relational control-rules. These rules are generated either (1) by reasoning from a firstorder domain description, or (2) more or less arbitrarily according to a taxonomic concept language. The cost of decision-theoretic planning in individual problems is substantial. State-of-the-art solution algorithms target either state-based (tabular) or factored propositional problem representations, thus they succumb to Bellman’s curse of dimensionality – i.e. the complexity of computing the optimal policy for a problem instance can be exponential in the dimension of the problem (Littman, Goldsmith, & Mundhenk 1998). A research direction which has garnered significant attention recently is that of generalisation in planning. The idea is that the cost of planning with propositional representations can be mitigated by technologies that plan for a domain rather than for individual problems. These approaches yield general policies which can be executed in any problem state from the domain at hand. In practice general policies are expressed in first-order/relational formalisms. Proposals to date suggest general policies can be achieved by either (1) reasoning