Value Iteration for Stochastic LQR With Convergence Guarantees

Jing Lai, Junlin Xiong, Yu Kang · IEEE Transactions on Neural Networks and Learning Systems · 2025

This brief studies the discounted stochastic linear quadratic regulator (LQR) problem for systems suffering from additive noise of unknown mean. A completely model-free (MF) value iteration (VI) algorithm is developed to learn the optimal control policy using off-line system trajectories. The generated control policies are proven to converge to a small neighborhood of the optimal ones with high probability. In addition, an MF algorithm is proposed to learn a feasible discount factor. The proposed MF algorithms are illustrated through several examples.

Read the paper · More papers on PaperTik