Sample Trajectory Selection Method Based on Large Language Model in Reinforcement Learning
Jianlin Lai, Zhaoxiang Zang · IEEE Access · 2024
This paper introduces a method for trajectory selection using large-scale pre-trained language models, aimed at improving sample efficiency and training efficiency in reinforcement learning. By utilizing a carefully designed prompt, we enable the large language model to fully utilize its prior knowledge, effectively understanding and assessing the quality of trajectories produced through agent-environment interactions in reinforcement learning. This approach allows selecting more informative trajectories for the current agent’s learning. Unlike other works that use large language models to indirectly improve reinforcement learning training efficiency by generating actions or decisions, our method employs these models to choose superior trajectories, thus more directly enhancing sample efficiency in reinforcement learning. We evaluated this approach on multiple benchmark tasks in OpenAI’s Gym and Rlcard. The results indicate a significant reduction in the number of environment interactions, with increasing the average reward by 37% compared to the original method.