A large language model-guided approach to the focused physical exam
Arya S Rao, Christian Rivera, Husayn F. Ramji, Sarah Wagner, Andrew Mu, John Kim, William Marks, Benjamin A. White, David C. Whitehead, Michael J. Senter-Zapata, Marc D. Succi · Journal of Medical Artificial Intelligence · 2025
Abstract: The physical exam is crucial in medical diagnosis, providing objective information that complements the patient’s history and guides clinical management. While studies show the potential of large language models (LLMs) like GPT-4 as adjunctive diagnostic tools, their use in physical exams remains unexplored. This study evaluates GPT-4’s ability to provide tailored physical exam instructions based on chief complaints. GPT-4 was prompted to recommend specific physical exam maneuvers for 19 chief complaints from the Hypothesis Driven Physical Exam Student Handbook by the American Association of Medical Colleges. Two board-certified emergency medicine and one internal medicine attending physicians evaluated GPT-4’s responses for accuracy, comprehensiveness, readability, and overall quality using a Likert scale (1= very poor, 3= neutral, 5= excellent) with subjective commentary. GPT-4 received average scores of 4.16 for accuracy, 3.95 for comprehensiveness, 4.39 for readability, and 3.89 for overall quality. The total average score per chief complaint was 49.16 out of 60. The highest score was for “leg pain upon exertion” [54], and the lowest was for “lower abdominal pain” [43]. The Pearson correlation coefficient showed a moderate association between overall quality and accuracy (r=0.61) and comprehensiveness (r=0.57). Reviewers highlighted GPT-4’s detailed and extensive outputs but noted the need to improve elements such as specificity and inclusion of vital signs. This study shows the potential of GPT-4 as an adjunctive diagnostic tool by providing useful physical exam recommendations. GPT-4 performed effectively in general medical scenarios, scoring at least 80% of the maximum possible points. Future research could compare unassisted physicians to those using LLM tools and explore custom-trained models, suggesting significant potential for LLMs in clinical decision support.