Extending Evolution-Guided Policy Gradient Learning into the multi-objective domain

Adam Callaghan, Karl Mason, Patrick Mannion · Neurocomputing · 2025

Multi-Objective Reinforcement Learning (MORL) poses significant challenges, primarily due to the necessity of balancing conflicting objectives—a limitation that traditional single-objective approaches fail to address. This paper introduces Multi-Objective Evolutionary Reinforcement Learning (MO-ERL), the first adaptation of Evolutionary Reinforcement Learning (ERL) specifically designed to address the complexities of the multi-objective domain effectively. MO-ERL integrates policy gradient-based reinforcement learning (RL), which optimizes expected utility, with evolutionary algorithms (EAs) that maintain diversity across the Pareto front. This combination leverages RL’s strength in exploitation and EAs’ proficiency in exploration, enabling MO-ERL to effectively navigate the trade-offs inherent in multi-objective optimization problems. Evaluation on multi-objective continuous control tasks using the MuJoCo physics engine demonstrates that MO-ERL outperforms state-of-the-art baselines, achieving up to 62.71% higher hypervolume and 196.28% greater expected utility. These results validate MO-ERL’s ability to balance solution diversity and optimality, setting a new benchmark for solving MORL tasks. • First extension of Evolutionary Reinforcement Learning to multi-objective domain. • Outperforms state-of-the-art (CAPQL, PCN) on multi-objective continuous control. • Up to 62.7% higher hypervolume, 196.3% utility on high-dimensional control tasks. • MO-GA hypervolume approximation reduces computational cost, maintaining performance. • Frequent MORL injections into MO-GA yield smoother learning and faster convergence.

Read the paper · More papers on PaperTik