Hierarchical Reinforcement Learning: Assignment of Behaviours to Subpolicies by Self-Organization

Wilco Moerman · 2009

“He knew what he had to do. It was, of course, an impossible task. But he was used to impossible tasks. (...) The way to deal with an impossible task was to chop it down into a number of merely very difficult tasks, and break each of them into a group of horribly hard tasks, and each one of them into tricky jobs, and each one of them...” Terry Pratchett — Truckers (Bromeliad Trilogy, book I) A new Hierarchical Reinforcement Learning algorithm called HABS (Hierarchical Assignment of Behaviours by Self-organizing) is proposed in this thesis. HABS uses self-organization to assign be-haviours to uncommitted subpolicies. Task decompositions are central in Hierarchical Reinforcement Learning, but in most approaches they need to be designed a priori, and the agent only needs to fill in the details in the fixed structure. In contrast, the new algorithm presented here autonomously identifies behaviours in an abstract higher level state space. Subpolicies self-organize to specialize for the high level behaviours that are actually needed. These subpolicies are then used as the high level actions.

Read the paper · More papers on PaperTik