Auto-Conditioned Stochastic Recursive Momentum for Policy Gradient Methods

Junhao Qin, Liu Liu, Yongchao Liu · 2025

Most existing policy gradient methods typically require knowledge of the problem-dependent parameters to set step sizes, which are usually unknown in many practical scenarios. In this paper, we propose an auto-conditioned stochastic recursive momentum policy gradient method (AC-STORMPG), which estimates local Lipschitz constants using first-order information from historical iterates and derives auto-conditioned step sizes from these estimates. We demonstrate that AC-STORMPG enjoys a sample complexity of O(ϵ −3) for finding an ϵ-approximate first-order stationary point, which matches the best-known rate for first-order policy gradient methods. Empirical results on classical reinforcement learning tasks verify the robustness and competitive performance of our algorithm.

Read the paper · More papers on PaperTik