On recursive temporal difference and eligibility traces

Simone Baldi, Di Liu, Zichen Zhang · IECON 2020 The 46th Annual Conference of the IEEE Industrial Electronics Society · 2020

This work studies a new reinforcement learning method in the framework of Recursive Least-Squares Temporal Difference (RLS-TD). Differently from the standard mechanism of eligibility traces, leading to RLS-TD(λ), in this work we show that the forgetting factor commonly used in gradient-based estimation has a similar role to the mechanism of eligibility traces. We adopt an instrumental variable perspective to illustrate this point and we propose a new algorithm, namely - RLS-TD with forgetting factor (RLS-TD-f). We test the proposed algorithm in a Policy Iteration setting, i.e. when the performance of an initially stabilizing controller must be improved. We take the cart-pole benchmark as experimental platform: extensive experiments show that the proposed RLS-TD algorithm exhibits larger performance improvements in the largest portion of the state space.

Read the paper · More papers on PaperTik