Auto-Conditioned Stochastic Recursive Momentum for Policy Gradient Methods
Junhao Qin, Liu Liu, Yongchao Liu · 2025
Most existing policy gradient methods typically require knowledge of the problem-dependent parameters to set step sizes, which are usually unknown in many practical scenarios. In this paper, we propose an auto-conditioned stochastic recursive momentum policy gradient method (AC-STORMPG), which estimates local Lipschitz constants using first-order information from historical iterates and derives auto-conditioned step sizes from these estimates. We demonstrate that AC-STORMPG enjoys a sample complexity of O(ϵ −3) for finding an ϵ-approximate first-order stationary point, which matches the best-known rate for first-order policy gradient methods. Empirical results on classical reinforcement learning tasks verify the robustness and competitive performance of our algorithm.