Automatic Policy Decomposition through Abstract State Space Dynamic Specialization

Rene Sturgeon, François Rivest · 2020

Significant progress has been made recently in deep reinforcement learning in the development of options. This idea consists in learning policies (or macro of actions) for sub-goals. An important bottleneck of this approach is that these options are often available as actions everywhere in the state space, hence, potentially enlarging the action space to search for the optimal policy. In this paper, we propose to use the fact that the state space is rarely fully connected, but instead has regions of highly connected states with fewer links between those regions. Our proposed model extends deep Q-Learning network (DQN) by splitting the top layers into multiple heads each specializing in learning the dynamics of a particular region of the state space as well as the optimal policy for that region. The state prediction quality of each head is used to determine which head is the local expert, rating its contribution to the current state's policy. We show that this approach is able to learn something similar to options and generalized value function, providing a promising alternative to the current approach.

Read the paper · More papers on PaperTik