Semi-supervised Fine-tuning for Large Language Models
Junyu Luo, Xiao Luo, Xiusi Chen, Zhiping Xiao, Wei Ju, Ming Zhang · 2025
Supervised fine-tuning (SFT) is crucial in adapting large language models (LLMs) to a specific domain or task.However, only a limited amount of labeled data is available in practical applications, which poses a severe challenge for SFT in yielding satisfactory results.Therefore, a data-efficient framework that can fully exploit labeled and unlabeled data for LLM fine-tuning is highly anticipated.Towards this end, we introduce a semi-supervised fine-tuning (SemiFT) task and a framework named SEMIEVOL for LLM alignment from a propagate-and-select manner.For knowledge propagation, SEMIEVOL adopts a bi-level approach, propagating knowledge from labeled data to unlabeled data through both inweight and in-context methods.For knowledge selection, SEMIEVOL incorporates a collaborative learning mechanism, selecting higherquality pseudo-response samples.We conducted experiments using GPT-4o-mini and Llama-3.1 on seven general or domain-specific datasets, demonstrating significant improvements in model performance on target data.Furthermore, we compared SEMIEVOL with SFT and self-evolution methods, highlighting its practicality in hybrid data scenarios.