An automated measure of MDP similarity for transfer in reinforcement learning
Haitham Bou Ammar, Eric R. Eaton, Matthew Edmund Taylor, Decebal Constantin Mocanu, Kurt Driessens, Gerhard Weiß, Karl Tuyls · 2014
Transfer learning can improve the reinforcement learn-ing of a new task by allowing the agent to reuse knowl-edge acquired from other source tasks. Despite their success, transfer learning methods rely on having rel-evant source tasks; transfer from inappropriate tasks can inhibit performance on the new task. For fully au-tonomous transfer, it is critical to have a method for automatically choosing relevant source tasks, which re-quires a similarity measure between Markov Decision Processes (MDPs). This issue has received little atten-tion, and is therefore still a largely open problem. This paper presents a data-driven automated simi-larity measure for MDPs. This novel measure is a sig-nificant step toward autonomous reinforcement learning transfer, allowing agents to: (1) characterize when trans-fer will be useful and, (2) automatically select tasks to use for transfer. The proposed measure is based on the reconstruction error of a restricted Boltzmann machine that attempts to model the behavioral dynamics of the two MDPs being compared. Empirical results illustrate that this measure is correlated with the performance of transfer and therefore can be used to identify similar source tasks for transfer learning.