RL with KL penalties is better viewed as Bayesian inference
Tomasz Korbak, Ethan Perez, Christopher L. Buckley · 2022
Reinforcement learning (RL) is frequently employed in fine-tuning large language models (LMs) to penalize them for undesirable features of generated sequences, such as offensiveness or harmfulness.In this paper, we analyze challenges associated with treating language models as RL policies and show how avoiding those challenges requires moving beyond the RL paradigm.We start by observing that naïve maximisation of the standard RL objective unavoidably leads to distribution collapse: turning the LM into a degenerate distribution.Then, we analyze KL-regularised RL, a widely used recipe for fine-tuning LMs, which additionally constrains the fine-tuned LM to stay close to its original distribution.We show that KL-regularised RL is equivalent to variational inference: approximating a Bayesian posterior which specifies how to update a prior LM to conform with evidence provided by the reward function.Based on these observations, we sketch a Bayesian perspective that provides a firstprinciples derivation for KL-regularised RL.The Bayesian perspective also separates the modelling problem (defining a target distribution specifying the desired behaviour of an LM) and the inference problem (approximating that target distribution).Finally, it suggests that RL is not a good formal framework for thinking about fine-tuning LMs.