Efficient Medical Question Answering Through QLoRA Fine-Tuning and Knowledge Distillation

Bartłomiej Brzęk, Grzegorz Dziczkowski · Procedia Computer Science · 2025

This work presents an efficient approach for fine-tuning and specialising large language models (LLMs) in the medical domain. The proposed method combines Quantised Low-Rank Adaptation (QLoRA) – a resource-efficient technique based on 4-bit quantisation and the update of a small set of parameters – with knowledge distillation (KD), which enables the transfer of specialised knowledge from a larger teacher model to a smaller student. As the teacher, a QLoRA-fine-tuned version of Meditron 7B is employed, while the student is Qwen2.5 3B – a significantly lighter model in terms of computational demands. This approach is evaluated on two medical question answering (QA) benchmarks: MedMCQA and MMLU-Medical. Results show that the smaller Qwen2.5 3B model, after undergoing KD, can reach accuracy levels comparable to – and in some cases even exceeding – those of its larger, domain-specialised teacher. Furthermore, compared to its baseline version, the student achieves an improvement of approximately 10 percentage points on both benchmarks. These experiments demonstrate that combining QLoRA with KD can effectively balance high-quality medical reasoning with reduced computational cost, making this approach promising for future applications of LLMs in resource-constrained clinical environments. While further validation and safety assessments are required before real-world deployment, the presented method offers a step toward more accessible and efficient domain-specific language models in healthcare.

Read the paper · More papers on PaperTik