Optimizing Response Consistency of Large Language Models in Medical Education through Prompt Engineering
Tharindu Weerathunge, Shantha Jayalal, Keerthi Wijayasiriwardhane · 2025
This research focuses on optimizing the response consistency of Large Language Models (LLMs) in medical education through advanced prompt engineering techniques. LLMs often give different answers to the same question, making self-consistency a critical parameter for assessing their performance. Addressing this inconsistency is essential in high-stakes fields like healthcare, where reliable and accurate information is important. The study employed custom prompt engineering strategies, including zero-shot, few-shot, and Chain-of-Thought (CoT) prompting, to improve LLM output consistency and accuracy. We implemented a retrieval-augmented generation (RAG) framework to use external knowledge from trusted medical resources, keeping the responses accurate and contextually appropriate. Responses were scored on several dimensions: content relevance, completeness, and clinical correctness, and assessed for consistency by asking repeated queries. The results showed significant enhancement in the consistency and accuracy of the responses, proving the effectiveness of the presented method. This work outlines suggestions for the use of LLMs in a way that can be incorporated into medical education while considering the limitations. It underscores the need for further exploration of prompt engineering to improve LLM performance and establishes these tools as reliable resources for training healthcare professionals.