Reward Design Using Large Language Models for Natural Language Explanation of Reinforcement Learning Agent Actions
Shinya Masadome, Taku Harada · IEEJ Transactions on Electrical and Electronic Engineering · 2025
Reinforcement learning (RL) has found applications across diverse domains; however, it grapples with challenges when formulating reward functions and exhibits low exploration efficiency. Recent studies leveraging large language models (LLMs) have made strides in addressing these issues. However, for RL agents to be practically deployable, elucidating their decision‐making process is crucial for enhancing explainability. We introduce a novel RL approach aimed at alleviating the burden of designing reward functions and facilitating natural language explanations for actions grounded in the agent's decisions. Our method employs two types of agents: a low‐level agent responsible for concrete action selection and a high‐level agent tasked with setting abstract action goals. The high‐level agent undergoes training using a hybrid reward function framework, which incentivizes its actions by comparing them with those generated by an LLM across discretized states. Meanwhile, the training of the low‐level agent is guided by a reward function designed using the EUREKA algorithm. We applied the proposed method to the cart‐pole problem and demonstrated its ability to achieve a learning convergence rate while reducing human effort. Moreover, our approach yields coherent natural language explanations elucidating the rationale behind the agent's actions. © 2025 The Author(s). IEEJ Transactions on Electrical and Electronic Engineering published by Institute of Electrical Engineers of Japan and Wiley Periodicals LLC.