FANNO: Augmenting High-Quality Instruction Data with Open-Sourced LLMs Only

He Zhu, Yifan Ding, Yicheng Tao, Zhiwen Ruan, Yixia Li, Wenjia Zhang, Yun Chen, Guanhua Chen · 2025

Instruction tuning stands as a crucial advancement in leveraging large language models (LLMs) for enhanced task performance.However, the annotation of instruction datasets has traditionally been expensive and laborious, often relying on manual annotations or costly proprietary LLMs.Recent works explore approaches to synthesize data with opensourced LLMs but require high-quality humancrafted seed data.In this work, we introduce FANNO, an end-to-end framework to synthesize high-quality instruction data with opensourced LLMs and sampled unlabeled documents, eliminating the necessity for seed data.Starting from diverse pre-screened documents, the framework synthesizes complex and diverse high-quality instruction and response pairs in different stages.We propose a tagging-based prompt method to generate diverse and complex seed data and a UCB-based approach to augment more instruction data with the seed data.A novel Think Different prompt is proposed to address the distributional limitations of the seeds, further boosting the data diversity.Experiments prove that the FANNO can generate diverse and complex high-quality data even with a opensource small teacher model.The synthesized instruction data demonstrates performance that is comparable to, or even surpasses, baseline annotation methods with proprietary LLMs or open-sourced LLMs while requiring fewer instruction data samples.

Read the paper · More papers on PaperTik