Prompt design for medical question answering with Large Language Models

Leonid Kuligin, Jacqueline Lammert, Aleksandr Ostapenko, Keno K. Bressem, Martin Boeker, Maximilian Tschochohei · Machine Learning with Applications · 2025

The combination of prompting technique and the choice of a foundational model determines end-to-end workflow performance on a given task. We aim to provide comprehensive guidance for the best-performing prompting techniques for various LLMs for medical question-answering. We aim to provide comprehensive guidance for the best-performing prompting techniques for a variety of LLM for medical question-answering. We evaluated 15 large LLMs (incl. Claude 3.5 Sonnet, Gemini pro, Llama, Mistral, OpenAI GPT-4o and 4.1) and 6 smaller models (incl. Gemma, Mistral Nemo, Llama 3.1, Gemini flash) across five prompting techniques on neuro-oncology exam questions. Using the established MedQA dataset and a novel neuro-oncology question set, we compared basic prompting, chain-of-thought reasoning, and more complex agent-based methods incorporating external search capabilities. Results showed that the Reasoning and Acting (ReAct) approach combined with giving LLM access to Google Search performed best on large models like Claude 3.5 Sonnet (81.7% accuracy and 85.5% for v2). We also showed that large models significantly outperformed smaller ones on the MedQA dataset (79.3% vs. 51.2% accuracy) and that complex agentic patterns like Language Agent Tree Search provided minimal benefits despite 5x higher latency. We recommend practitioners to experiment with various techniques given their specific use case and foundational model, and favor simple prompting patterns with large models, as they offer the best balance of accuracy and efficiency.

Read the paper · More papers on PaperTik