Q-learning with Uniformly Bounded Variance: Large Discounting is Not a Barrier to Fast Learning
Adithya M. Devraj, Sean Meyn · arXiv (Cornell University) · 2020
Sample complexity bounds are a common performance metric in the Reinforcement Learning literature. In the discounted cost, infinite horizon setting, all of the known bounds have a factor that is a polynomial in $1/(1-γ)$, where $γ 0$ is an upper bound on the spectral gap of an optimal transition matrix.