Reducing Diagnostic Uncertainty in Emergency Departments: The Role of Large Language Models in Age-Specific Diagnostics
Wanting Cui, Kensaku Kawamoto, Keaton L. Morgan, Joseph Finkelstein · 2024
Diagnostic errors in emergency departments affect approximately 7.4 million patients every year. To address this, the integration of artificial intelligence, specifically Large Language Models (LLMs) like ClinicalBERT, into differential diagnostic process has been explored to reduce the diagnostic uncertainty and alleviate physician workload. This study focused on assessing the variance in diagnostic accuracy of LLMs between young and middle-aged adults (18–64 years) and older adults (65+ years) using 13 models fine-tuned on the MIMIC-III dataset, each targeting a specific body system. There were 8,321 cases of hospital stays and 124,736 clinical notes in the analytic dataset. Results indicated that while some models performed consistently well across both age groups, there was a discernible variability in others. 77% of models showed strong performance for younger adults, compared to 54%for older adults. The neoplasm (NEO) and circulatory (CIR) models stood out in both groups. In addition, the respiratory systems showed better performance for older adults. The mental and behavioral (MBD) system models, however, demonstrated a significant decline in recall for older adults. These findings highlight the importance of incorporating age-specific adjustments into AI models to optimize diagnostic precision across diverse patient populations.