Transfer of task representation in reinforcement learning using policy-based proto-value functions
Eliseo Ferrante, Alessandro Lazaric, Marcello Restelli · 2008
Reinforcement Learning research is traditionally devoted to solve single-task problems. Therefore, anytime a new task is faced, learning must be restarted from scratch. Recently, several studies have addressed the issue of reusing the knowl-edge acquired in solving previous related tasks by transfer-ring information about policies and value functions. In this paper, we analyze the use of proto-value functions under the transfer learning perspective. Proto-value functions are effective basis functions for the approximation of value func-tions defined over the graph obtained by a random walk on the environment. The definition of this graph is a key as-pect in transfer transfer problems in which both the reward function and the dynamics change. Therefore, we introduce policy-based proto-value functions, which can be obtained by considering the graph generated by a random walk guided by the optimal policy of one of the tasks at hand. We compare the effectiveness of policy-based and standard proto-value functions, on different transfer problems defined on a simple grid-world environment.