Real-time Integration of Fine-tuned Large Language Model for Improved Decision-Making in Reinforcement Learning
Xiancai Xiang, Jian Yang Xue, Lin Zhao, Yuan Lei, Chao Yue, Ke Lü · 2024
In this paper, we investigate a novel and efficacious methodology for training reinforcement learning agents. We also further reveal the potential of the Large Language Model (LLM) in intricate decision-making environments. Reinforcement learning, one main approach in training decision-making capabilities for intelligent agents, suffers from the challenges of sparse rewards and inefficient exploration. Considering the current phenomenal performance of pre-trained LLMs, certain studies have adopted the LLMs to shepherd the actions of intelligent agents. However, there are some limitations in the application of LLMs, such as the lack of domain knowledge of LLMs for specific fields, and LLMs generally access a slower response time but consume a higher economic cost. Embarking from these constraints, this paper takes on the formidable challenge of the Unmanned Air Vehicle (UAV) air combat simulations environment, where decision-making is notably circumscribed by temporal limitations. We initially fine-tuned the LLMs to a domain-specific expertise while concurrently constructing a knowledge base. Hence, upon gaining profound insight into the field of air combat, it evolves into an adept model capable of effectively guiding UAV decisions. Further addressing the difficulties of real-time accessing the LLM, we propose a novel approach named Reward Shaping with Large Language Model (LLM-RS) to augment the autonomous decision-making competency of UAVs within the context of air combat simulations. We compared the agents trained by conventional reinforcement learning techniques and those trained by LLM-RS. The experimental results reveal that, under the same equipment training conditions, the LLM-RS technique grounded in the fine-tuning of the LLM and the knowledge base substantially improves the performance of UAVs in air combat simulations, while concurrently diminishing the requisite training duration.