Task-Specific Bidirectional Knowledge Distillation: Transforming General Large Language Models Into Lightweight Models
Zeying Jin, Yujie Wang · 2025
Large language models have shown impressive capabilities in natural language processing tasks, but their deployment is hindered by high computational costs and lack of domain specificity. Existing knowledge distillation methods, which transfer knowledge from large teacher models to smaller student models, often fail to capture the nuances of specialized tasks owing to domain differences. To address this issue, we propose a task-specific bidirectional knowledge distillation (TBKD) method. This method fine-tunes the teacher model on domain-specific data to capture key features, and then distills it into a student model using bidirectional KL divergence, balancing coverage, and diversity. Experiments on a petrochemical fire prevention dataset show that our method achieves lower perplexity and higher BLEU scores than existing methods, demonstrating its effectiveness in compressing large language models while preserving task-specific performance.