Rethinking Task-Specific Knowledge Distillation: Contextualized Corpus as Better Textbook
Chang Liu, Chongyang Tao, Jianxin Liang, Tao Shen, Jiazhan Feng, Quzhe Huang, Dongyan Zhao · 2022
Knowledge distillation has been proven effective when customizing small language models for specific tasks.Here, a corpus as 'textbook' plays an indispensable role, only through which the teacher can teach the student.Prevailing methods adopt a two-stage distillation paradigm: general distillation first with taskagnostic general corpus and task-specific distillation next with augmented task-specific corpus.We argue that such a paradigm may not be optimal.In general distillation, it's extravagant to let the diverse but desultory general knowledge overwhelms the limited model capacity of the student.While in task-specific distillation, the task corpus is usually limited and narrow, preventing the student from learning enough knowledge.To mitigate the issues in the two gapped corpora, we present a better textbook for the student to learn: contextualized corpus that contextualizes task corpus with large-scale general corpus through relevance-based text retrieval.Experimental results on GLUE benchmark demonstrate that contextualized corpus is the better textbook compared with jointly using general corpus and augmented task-specific corpus.Surprisingly, it enables task-specific distillation from scratch without general distillation while maintaining comparable performance, making it more flexible to customize the student model with desired model size under various computation constraints.