A Survey on Quality Evaluation of Instruction Fine-tuning Datasets for Large Language Models

Yitian Luo, Yu Liu, Lu Zhang, Feng Gao, Siyi Gu · Data Intelligence · 2025

Instruction fine-tuning is a key method for adapting large language models (LLMs) to domain-specific tasks, and instruction quality significantly impacts model performance after fine-tuning. Hence, evaluating the quality of instruction and selecting high-quality instructions are essential steps in the process of LLM instruction fine-tuning. Although existing studies provide important theoretical foundations and techniques for this, there is still room for improvement in terms of generality, the relationship between methods and experimental verification. Current methods for evaluating instruction quality can be classified into four main categories: human evaluation, statistics-based evaluation, model-based evaluation, and LLMs-based evaluation. Among these methods, human evaluation relies on the subjective judgment and domain expertise of the evaluators, which offers interpretability and is suitable for scenarios involving small-scale data and sufficient budgets. Statistics-based evaluation estimates the quality of instructions using indicators such as stopwords and lexical diversity, providing high efficiency and a suitable evaluation for large-scale data. Model-based evaluation employs specific models to quantify indicators such as perplexity (PPL) and instruction following difficulty (IFD), which is flexible and suitable for specific tasks. The LLMs-based evaluation rates the quality of instructions through prompt-based interaction with LLMs, focusing on aspects such as accuracy and coherence, which is highly automated and customizable, simplifying the evaluation process. Finally, considering the limitations of current quality evaluation methods, some future research directions are proposed for improvement. These include refining instruction categories, extending evaluation indicators, enhancing human-AI interaction evaluation method, applying agents in instruction quality evaluation, and developing a comprehensive evaluation framework.

Read the paper · More papers on PaperTik