Regularized Q-Learning With Linear Function Approximation
Jiachen Xi, Alfredo Garcia, Petar Momčilović · IEEE Transactions on Automatic Control · 2025
We consider a single-loop algorithm for regularized Q-learning with linear function approximation. The proposed algorithm is motivated by a bi-level optimization formulation of regularized Q-learning wherein thelowerlevel optimization problem aims to identify a value function approximation that satisfies Bellman's recursive optimality condition and theupperlevel aims to find the projection onto the span of basis vectors. We show that, under certain assumptions, the proposed algorithm converges to a stationary point in the presence of Markovian noise. In addition, we provide a performance guarantee for the policies derived from the proposed algorithm.