Policy Transfer via Skill Adaptation and Composition
Benhui Zhuang, Chunhong Zhang, Zheng Hu · 2022
While reinforcement learning is generally powerful on decision making, it usually suffers from slow convergence rate on complex problems. Enlightened by human experience, transfer learning that reuses skills learned from the source task to a new target task is desirable to speed up the learning process. However, most existing works achieve policy transfer by reusing fixed skills learned from the source task and simply recomposing them linearly in the target task, which is inefficient to adapt skills to the differences between source and target tasks. To address this issue, we propose a framework, Pre-Trained Hierarchical Recurrent Independent Mechanisms (PHRIM), that optimizes a hierarchical policy for selective composition of modular skills and learns adaptive skills in addition to the transferred skills for the adaptation of task differences. The experiment conducted on MiniGrid maze game demonstrates that the PHRIM outperforms the fixed-skill baselines on adapting to the source-to-target differences from aspects of state, action, and task complexity respectively.