Evolutionary Policy Optimization
Zelal Su Mustafaoglu, Keshav K. Pingali, Risto P Miikkulainen · Proceedings of the Genetic and Evolutionary Computation Conference Companion · 2025
A key challenge in reinforcement learning (RL) is managing the exploration-exploitation trade-off without sacrificing sample efficiency. Policy gradient (PG) methods excel in exploitation through fine-grained, gradient-based optimization but often struggle with exploration due to their focus on local search. In contrast, evolutionary computation (EC) methods excel in global exploration. To address these limitations, this paper proposes Evolutionary Policy Optimization (EPO), a hybrid algorithm that integrates neuroevolution with policy gradient methods for policy optimization. EPO leverages the exploration capabilities of EC and the exploitation strengths of PG, offering an efficient solution to the exploration-exploitation dilemma in RL. Experiments with the Atari Pong and Breakout benchmarks show that EPO improves both policy quality and sample efficiency compared to standard PG and EC methods.