Training Large Language Models to Follow System Prompt with Self-Supervised Fine-tuning

Junyan Qiu, Yiping Yang · 2024

In the realm of artificial intelligence, system prompts stand as directives or requests aimed at guiding systems, such as programming environments or AI models, to execute specific tasks or operations. Typically positioned at the commencement of input sequences in large language models, these prompts play a pivotal role in shaping the model’s response and guiding its interaction flow. However, a notable challenge emerges during multi-turn dialogues, where these models gradually diverge from adhering to the initial system prompt, leading to inconsistencies in the dialogue. In this paper, we present a scalable framework facilitating the adherence of language models to system prompts through automated data construction. Our approach, termed Self-Supervised System Prompt Fine-tuning (S3FT), begins by prompting a language model to modify real dialogue responses to fit a specific system prompt, using stylized translation. Subsequently, we select a small sample of these responses for human preference annotation. This annotated data is utilized to train the language model to act as a discriminator, identifying high-quality examples that are then employed in further supervised fine-tuning. Experimental results on several datasets demonstrate that applying our method to LlaMA2 and ChatGLM promotes human preference rates by over 50%, and outperforms ChatGPT and GPT4 by a consideratble margin.

Read the paper · More papers on PaperTik