Leveraging Large Language Models for Creating Clinical Decision Support for Pneumonia Severity Triage
Jonathan R. Tiao, Kyle Kulas, D.S. Videlefsky, Ying Kuen Cheung, Máire O'Donnell, Benjamin L. Ranard · American Journal of Respiratory and Critical Care Medicine · 2025
Abstract RATIONALE Accurate prediction of pneumonia severity ensures that patients presenting to the Emergency Department are triaged to the appropriate level of care. While there have been many attempts to create clinical decision support tools, these attempts have largely relied on variables such as demographics, vitals, and laboratory data, with limited options to leverage data such as notes and imaging. Advances in machine learning, most notably large language models (LLMs), allow for new approaches to these types of data to train prediction models. Here we demonstrate the utilization of a LLM as a feature to predict severity of pneumonia. METHODS Vitals, laboratory, demographic data, and chest radiographs were obtained on patients admitted to a single center in New York City from 1/2017 – 9/2019 with an ICD-10 primary diagnosis of pneumonia present on admission (J9-J18). The highest reported respiratory rate, highest reported temperature, lowest systolic blood pressure, and highest pulse within the first 24 hours of presentation were utilized. The first set of labs on arrival were used. The first chest radiograph report from each presentation was then pre-processed to remove names and dates. Llama-3, an open source LLM, was then applied to the reports to identify if the radiology report noted whether there was evidence of a consolidation, if there were findings consistent with atypical pneumonia, and how many sides of the lungs were affected by pneumonia. In total 17 variables were used – 4 vital signs, 10 labs, and 4 radiograph-based variables. The outcome variable for the multivariate logistic regression was whether the patient would require continuous bilevel, high-flow nasal cannula, or intubation at any point during their admission. Samples with missing data were dropped. Significance was set as a p-value < 0.05. Modeling was performed in python 3, Llama-3 was run locally using Ollama. RESULTS Complete data were available for 2,413 samples. 284 samples required higher levels of respiratory support. Results of the multivariable regression model are illustrated in Figure 1. The radiographic involvement of pneumonia on both sides was statistically significant (p = 0.02) as was the presence of atypical pneumonia (p=0.007). CONCLUSIONS LLMs present a novel way to process language for creation of clinical decision support. These findings suggest that it is feasible to use LLMs to create features from radiology reports in clinical decision support prediction tools. Further work is needed to evaluate if these features can improve existing pneumonia severity triage tools. Figure 1