Leveraging LLMs for Early Diagnosis in the Emergency Department: Comparing ClinicalBERT and GPT-4
Wanting Cui, Joseph Finkelstein · Studies in health technology and informatics · 2025
This study explored the potential of LLMs, such as ClinicalBERT and GPT-4, to identify potential diagnoses using early clinical notes from the MIMIC-III dataset. We compared these models across four conditions: circulatory system diseases, respiratory system diseases, septicemia, and pneumonia. ClinicalBERT consistently outperformed the GPT models, with its highest F1-score of 0.952 for respiratory system diseases. The GPT models, while showing high recall, had lower precision, with the highest F1-score of 0.784 achieved by the GPT binary voting method. ClinicalBERT demonstrated strong precision and F1-scores, while GPT-4 excelled in recall.