Evolutionary Reinforcement Learning with LLM-Based Hyperparameter Adaptation
Beining Chen, Feng-Feng Wei, Wei–Neng Chen · 2025
Evolutionary Reinforcement Learning (ERL), combining the exploration capability of evolutionary algorithms with reinforcement learning, has achieved remarkable success in various decision-making tasks. However, its performance is highly sensitive to the selection of hyperparameters. In recent years, the application of Large Language Models (LLMs) in automated optimization tasks has shown tremendous potential. Based on this, this study explores the role of LLMs in the hyperparameter tuning of ERL, which is named as LLM-optimized ERL. LLMoptimized ERL adopts Alibaba's Qwen-max model during the training process, connecting to the Qwen large model via the OpenAI API client. During invocation, a request is constructed containing the fitness values from different generations in the current training process, asking the Qwen model to generate recommendations for adjusting hyperparameters$\alpha$(learning rate) and$\gamma$(discount factor) based on this information. Comparing the performance with DRL and ERL across different map sizes$(4 \times 4,6 \times 6,8 \times 8,16 \times 16)$within the FrozenLake-v1 environment, the experimental results indicate that LLM-optimized ERL can dynamically adjust hyperparameters$\alpha$and$\gamma$according to environmental dynamics. This paper verifies the feasibility of LLMs in the domain of ERL and reveals their potential value in policy optimization and automatic tuning.