On the Feasibility of Language Modelling via Reinforcement Learning in Embedding Space

Yongchao Huang · INRIA a CCSD electronic archive server · 2025

Standard language models are trained via supervised learning, using cross-entropy loss to maximize the probability of the correct next token. This paper explores a novel alternative: framing next-word prediction as a reinforcement learning (RL) problem in a continuous embedding space. We propose a method where an Actor-Critic agent learns to regress directly on word embeddings, taking the preceding text as state and predicting the embedding vector of the next word as a continuous action. The reward is defined by the cosine similarity between the predicted and true embeddings. Through a progressive four-stage case study, we demonstrate the initial feasibility of this approach on a toy dataset and systematically address challenges such as training instability by incorporating techniques such as Target Networks. However, when scaled to a natural text corpus, the model's performance is ultimately hindered by the inherent sample inefficiency of RL. The agent falls into mode collapse, learning a suboptimal policy of predicting common, low-risk words. We conclude that while theoretically feasible, this RL-based regression approach is not practically competitive with traditional supervised methods for language modelling due to the sparse and less informative nature of the RL reward signal compared to direct gradient-based supervision.

Read the paper · More papers on PaperTik