Evaluating the accuracy of large language models in pulmonary medicine scientific writing
Sana Dang, Michelle Melfi, Phillip J. Gary · Journal of Medical Artificial Intelligence · 2025
Background: Large language models (LLMs) have vast potential applications in healthcare across the clinical, educational, and research settings. The tendency of these models to fabricate plausible sounding but factually incorrect information is referred to as artificial hallucination and is a significant limitation to their use. Studies evaluating the accuracy of these models in scientific writing across various subspecialties have demonstrated a variable yet alarming frequency of hallucinations in LLM generated output. This study aimed to evaluate the accuracy of LLMs in scientific writing specific to the field of pulmonary medicine. Methods: Five commonly encountered conditions relevant to pulmonary medicine were selected: asthma, chronic obstructive pulmonary disease, interstitial lung disease, pneumonia, and lung adenocarcinoma. A new artificial intelligence (AI) chatbot in Google Gemini was tasked with providing ten research articles for each condition with complete citations, which were then appraised by comparing them to the original articles for accuracy by two independent reviewers using the search tool on Google and PubMed for verification. Results: A high degree of agreement was observed between reviewers (κ=0.9277). Of the 50 total citations, only one citation (2%) relevant to pneumonia was found to be completely accurate across all citation components. Article title was found to have the highest accuracy with thirty correct results (60%) while DOI performed the poorest with only two correct results (4%). Conclusions: Only 2% of the citations generated by Google Gemini were found to be completely accurate across all citation components. A high degree of fabrication of plausible sounding but factually incorrect results, or artificial hallucinations, was observed. Thus, continued vigilance and rigorous efforts are necessary to validate the authenticity of responses generated by LLMs, especially as pertains to medical literature.