Enhancing Medical Diagnostic Reasoning with Chain-of-Thought in Large Language Models
Guangxin Dai, Xiang Li, Lizhou Fan, Xin Ma · 2025
To enhance the diagnostic reasoning capabilities of Large Language Models (LLMs) in clinical settings, the paper proposes four Chain-of-Thought (CoT) prompting strategies: Diagnosis Validation, Bayesian Diagnosis, Hierarchical Reasoning, and Causal Abduction. The CoT is designed to emulate distinct modes of clinical reasoning. Using GPT-3.5 and GPT-4o, we filtered the MedQA-USMLE dataset to identify 119 consistently error-prone questions from an initial pool of 1597. Experimental evaluation on GPT-40 shows that the Causal Abduction CoT achieves the highest accuracy ($35.3 \%$), followed by Diagnosis Validation (31.1%), Hierarchical Reasoning (26.9%), and Bayesian Diagnosis ($\mathbf{2 0. 2 \%}$). The results demonstrate that clinical reasoning-oriented CoT can effectively improve LLM performance on challenging diagnostic tasks.