Improving the Contextual Understanding of LLMs Through Multi-Teacher Knowledge Distillation and RAG
Yashaswini Ippili, Vruddhi K Jain, Yash Dixit, Gautam Saraf, V R Badri Prasad · 2025
Large-scale language models (LLMs) excel in various NLP tasks but face challenges in resource-constrained environments due to high computational and memory demands. To overcome this, we propose an architecture combining multiteacher knowledge distillation (MTKD) and retrieval-augmented generation (RAG) to improve the performance of smaller, efficient models without sacrificing accuracy. By using multiple teacher models, we transfer diverse knowledge to the student model, preserving its ability to handle complex tasks. RAG further enhances accuracy by dynamically retrieving relevant context during inference. Our experiments show that this approach outperforms traditional distillation methods, providing precise, context-aware responses while maintaining computational efficiency. The model is optimized for deployment on consumergrade GPUs.