Convergence of Q-learning with linear function approximation

Francisco S. Melo, Isabel Ribeiro · 2007

In this paper, we analyze the convergence properties of Q-learning using linear function approximation. This algorithm can be seen as an extension to stochastic control settings of TD-learning using linear function approximation, as described in [1]. We derive a set of conditions that implies the convergence of this approximation method with probability 1, when a fixed learning policy is used. We provide an interpretation of the obtained approximation as a fixed point of a Bellman-like operator. We then discuss the relation of our result with several related works as well as its general applicability.

Read the paper · More papers on PaperTik