SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Yi Dong, Zhilin Wang, Makesh Narsimhan Sreedhar, Xianchao Wu, Oleksii Kuchaiev · 2023

Model alignment with human preferences is an essential step in making Large Language Models (LLMs) helpful and consistent with human values.It typically consists of supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) stages.However, RLHF faces inherent limitations stemming from a complex training setup and its tendency to align the model with implicit values that end users cannot control at run-time.Moreover, reward models in RLHF stage commonly rely on single-dimensional feedback as opposed to explicit, multifaceted signals that indicate attributes such as helpfulness, humor, and toxicity.To address these limitations, we propose STEERLM, a supervised finetuning method that empowers end-users to control responses during inference.STEERLM conditions responses to conform to an explicitly defined multi-dimensional set of attributes, thereby empowering a steerable AI capable of generating helpful and high-quality responses while maintaining customizability.Experiments show that STEERLM trained on open source datasets generates responses that are preferred by human and automatic evaluators to many state-of-the-art baselines trained with RLHF while being much easier to train.Try STEERLM at https://huggingface.co/ nvidia/SteerLM-llama2-13B

Read the paper · More papers on PaperTik