Least-Squares SARSA(Lambda) Algorithms for Reinforcement Learning
Shenglei Chen, Yan-Mei Wei · 2008
The problem of slow convergence speed and low efficiency of experience exploitation in SARSA(lambda) learning is analyzed. And then the least-squares approximation model of the state-action pair's value function is constructed according to current and previous experiences. A set of linear equations is derived, which is satisfied by the weight vector of function approximator on a set of basis. Thus the fast and practical least-squares SARSA(lambda) algorithm and improved recursive algorithm are proposed. The experiment of inverted pendulum demonstrates that these algorithms can effectively improve convergence speed and the efficiency of experience exploitation.