Enhancing Clinical Decision Support with LLMs: A Feasibility Study Integrating CoT, RAG, and QLoRA
George Matthew, Yulia Hicks · Procedia Computer Science · 2025
The integration of artificial intelligence (AI) into healthcare holds significant promise for enhancing clinical decision-making and improving patient outcomes. However, general-purpose large language models (LLMs) frequently exhibit limitations such as hallucinations, lack of domain-specific accuracy, and opaque reasoning processes, posing risks in clinical applications. This study addressed these challenges by proposing and exploring an innovative integration of Chain-of-Thought (CoT) prompting, Retrieval-Augmented Generation (RAG), and parameter-efficient fine-tuning using Quantized Low-Rank Adaptation (QLoRA) specifically tailored for medical use. A distilled 14-billion-parameter variant of the DeepSeek R1 model was fine-tuned using a structured clinical dataset that emphasizes step-by-step reasoning. Additionally, external medical references were incorporated through RAG, employing embedding models for precise context retrieval. The combination of these techniques was systematically evaluated on a challenging set of open-ended medical questions, resulting in accuracy improvements—from a baseline accuracy of 55% to a final performance of 81%. Further qualitative evaluation involving three practising General Practitioners (GPs) and three fourth-year medical students from Cardiff University underscored the proposed system’s clinical utility and transparent reasoning capabilities while also identifying areas for improvement, such as conciseness and explicit adherence to national-specific clinical guidelines. This research demonstrated that integrating CoT, RAG, and QLoRA provides a practical pathway toward reliable, transparent, and clinically relevant AI support for healthcare professionals. Recommendations for future work include scaling models, incorporating comprehensive patient data, and enhancing customization for clinical application contexts.