Use of a large language model (LLM) for pan-cancer automated detection of anti-cancer therapy toxicities and translational toxicity research.

Ziad Bakouny, Nishat Ahmed, Chris Fong, Afsana Rahman, Tomin Perea-Chamblee, Karl Pichotta, Michele Waters, Chenlian Fu, Mark Yungjie Jeng, Mindy Lee, Chris Marotta, Neil J. Shah, Nikolaus D. Schultz, Adam Jacob Schoenfeld, Robert John Motzer, Justin Jee, Craig B. Thompson, Jian Carrot‐Zhang, Eduard Reznik, MSK Cancer Data Science Initiative (CDSI) · Journal of Clinical Oncology · 2025

1558 Background: Understanding why patients develop adverse events to anti-cancer therapies and predicting the occurrence of these toxicities has lagged behind tumor response biomarker development. This critical gap is primarily due to limited availability of large-scale curated toxicity data. Here, we leverage advances in natural language processing ( Jee J et al., Nature, 2024 ), pooled clinical trial data, and associated germline sequencing to detect adverse event data and determine clinical and genomic correlates. Methods: We utilized the Llama 3.1 LLM to automatically annotate patient adverse event data for 5 of the most common anti-cancer therapy related adverse events (adrenal insufficiency, hyperthyroidism, hypothyroidism, colitis, and pneumonitis). To validate LLM predictions at the patient-level, we used a pooled institutional dataset with gold standard prospectively collected adverse event data from 1,754 patients with solid tumors across 675 individual clinical trials. We further validated the LLM predictions at the clinical note-level using a subset of 100 manually curated notes. We evaluated note-level and patient-level predictions using sensitivity and specificity. Patient-level time-to-adverse event development predictions were evaluated using Pearson R 2 coefficients. Common Terminology Criteria for Adverse Events v 5.0 was used for toxicity definitions. Results: The patients’ average age (standard deviation) was 61.6 (14.5) years and 836 (47.7%) were female. The most common cancers were non-small cell lung cancer (N= 194, 11.1%), soft tissue sarcoma (N=171, 9.7%), breast cancer (N= 155, 8.8%), and melanoma (N=129, 7.4%). 44 (2.5%) patients had adrenal insufficiency, 88 colitis (5.0%), 253 hypothyroidism (14.4%), 66 hyperthyroidism (4.4%), and 146 pneumonitis (8.3%). Among 1258 patients with complete systemic therapy information available, 422 (33.5%) were treated with immunotherapy and 563 (44.8%) with chemotherapy. The performance metrics for LLM predictions at the note and patient levels are summarized in the table. Conclusions: We demonstrate the ability of an LLM to accurately annotate anti-cancer therapy toxicity data across a large number of patients. This approach is scalable to other toxicities and promises to spur adverse event research. Clinical and genomic correlates of anti-cancer therapy adverse events, using data from all patients with solid tumors with MSK-IMPACT data, will also be presented at the meeting. Performance metrics for LLM model. Toxicity Note-level (N= 100 notes) Patient-level (N= 1,754 patients) Sensitivity Specificity Sensitivity Specificity R 2 Adrenal insufficiency 100.0% 97.8% 97.7% 94.7% 98.2% Colitis 66.7% 99.0% 94.3% 80.4% 89.2% Hyperthyroidism 57.1% 100.0% 74.0% 91.4% 98.7% Hypothyroidism 100.0% 88.9% 88.1% 74.0% 96.1% Pneumonitis 76.9% 97.7% 98.6% 70.1% 83.9%

Read the paper · More papers on PaperTik