Enhancing GPT-2 Text Generation Through Reinforcement Learning and Token-Based Environments

Kanika Debbarma, Koj Sambyo · 2025

This study investigates the use of the Gym library to combine reinforcement learning with a pre-trained language model, GPT-2 to generate text in a customised setting. Using reward-basedlearning, the model is trained to optimise text creation. The environment's goal is to produce text from user input while adhering to distinct actions that represent token from the GPT-2 tokeniser. Throughout training, the model is prodded to get better and better inside this random reward system. We describe in detail how to use the Adam optimiser and train and test the model in this context. We also put up a testing mechanism that will yield performance metrics on text production through pre-programmed prompts. The model's output from the experiment shows highly repetitive behaviour, which highlights the need for more intricate reward systems and exploration strategies. Subsequent future work will focus on refining the reward function, utilising advanced exploration strategies and expanding the approach to larger language models.

Read the paper · More papers on PaperTik