On the Locality of Action Domination in Sequential Decision Making
Emmanuel Rachelson, Michail G. Lagoudakis · 2010
In the field of sequential decision making and reinforcement learning, it has been observed that good policies for most problems exhibit a significant amount of structure. In prac-tice, this implies that when a learning agent discovers an ac-tion is better than any other in a given state, this action ac-tually happens to also dominate in a certain neighbourhood around that state. This paper presents new results proving that this notion of locality in action domination can be linked to the smoothness of the environment’s underlying stochastic model. Namely, we link the Lipschitz continuity of a Markov Decision Process to the Lispchitz continuity of its policies’ value functions and introduce the key concept of influence ra-dius to describe the neighbourhood of states where the dom-inating action is guaranteed to be constant. These ideas are directly exploited into the proposed Localized Policy Itera-tion (LPI) algorithm, which is an active learning version of Rollout-based Policy Iteration. Preliminary results on the In-verted Pendulum domain demonstrate the viability and the potential of the proposed approach. 1