Enhancing substance use detection in clinical notes with large language models
Fabrice Harel-Canada, Anabel Salimian, Brandon Moghanian, Sarah E. Clingan, Allan Nguyen, Tucker D. Avra, Michelle Poimboeuf, Ruby Romero, Arthur Funnell, Panayiotis Petousis, Michael Shin, Nanyun Peng, Chelsea L. Shover, David Goodman‐Meza · Drug and Alcohol Dependence · 2025
Identifying substance use behaviors in electronic health records (EHRs) is challenging because critical details are often buried in unstructured notes that use varied terminology and negation, requiring careful contextual interpretation to distinguish relevant use from historical mentions or denials. Using MIMIC-III/IV discharge summaries, we created a large, annotated drug detection dataset to tackle this problem and support future systemic substance use surveillance. We then investigated the performance of multiple large language models (LLMs) for detecting eight substance use categories within this data. Evaluating models in zero-shot, few-shot, and fine-tuning configurations, we found that a fine-tuned model, Llama-DrugDetector-70B, outperformed others. It achieved near-perfect F1-scores ( ≥ 0 . 95 ) for most individual substances and strong scores for more complex tasks like prescription opioid misuse (F1=0.815) and polysubstance use (F1=0.917). These findings demonstrated that LLMs significantly enhance detection, showing promise for clinical decision support and research, although further work on scalability is warranted. • Large language models (LLMs) show high accuracy for detecting substance use in clinical notes ( ⩾ 0.95). • Fine-tuned open-source LLMs outperform proprietary tools for substance use detection. • New annotated dataset was released to advance substance use detection research.