Incremental least squares policy iteration in reinforcement learning for control
Chungui Li, Meng Wang, Shuhong Yang · 2008
We propose a novel algorithm of reinforcement learning for control problems which combines value-function approximation with linear architectures and approximate policy iteration. This algorithm improves least-squares policy iteration (LSPI) methods by using incremental least-squares temporal-difference learning algorithm (iLSTD) for prediction problems. We show that the novel algorithm has less computing complexities than LSPI, and has the same performance as LSPI in learning optimal policies.