LLM-Based Reward Engineering for Reinforcement Learning: A Chain of Thought Approach
Xinning Zhu, Jinxin Du, Qiongying Fu, Lunde Chen · 2025
Reinforcement Learning (RL) has achieved significant milestones in various domains. Traditional reward engineering often requires extensive domain knowledge and trial-and-error experimentation. This paper explores the integration of Large Language Models (LLMs) with CoT reasoning to automate and enhance the generation of reward functions for RL environments. We present a comprehensive methodology that leverages CoT-enabled LLMs to generate sophisticated, adaptive, and context-aware reward functions, implemented within Gymnasium environments and trained using Stable Baselines3. Through a detailed case study on the Bipedal Walker environment, we demonstrate the efficacy of our approach in producing superior agent performance compared to non-CoT approaches.