Incremental least squares policy iteration in reinforcement learning for control

Chungui Li, Meng Wang, Shuhong Yang · 2008

We propose a novel algorithm of reinforcement learning for control problems which combines value-function approximation with linear architectures and approximate policy iteration. This algorithm improves least-squares policy iteration (LSPI) methods by using incremental least-squares temporal-difference learning algorithm (iLSTD) for prediction problems. We show that the novel algorithm has less computing complexities than LSPI, and has the same performance as LSPI in learning optimal policies.

Read the paper · More papers on PaperTik