The Asymptotic Convergence-Rate of Q-learning
Csaba Szepesvári · 1997
In this paper we show that for discounted MDPs with discount factor> 1=2 the asymptotic rate of convergence of Q-learning is O(1=t R(1 �)) if R(1 � ) 0, where pmin and pmax now become the minimum and maximum state-action occupation frequencies corresponding to the stationary distribution. 1