Inferring ECOG performance status (PS) using large language models in patients with advanced prostate cancer.

Ammad Raina, Miguel Muniz, Muhammad Umair Anjum, Umair Ayub, Syed Arsalan Ahmed Naqvi, Salman Ayub Jajja, Ji-Eun Irene Yum, Ben Zhi Zhou, Nathan Y. Yu, Haidar M. Abdul‐Muhsin, Alton Oliver Sartor, Jacob J. Orme, Parminder Pal Singh, Yousef Zakharia, Daniel S. Childs, Irbaz Bin Riaz · Journal of Clinical Oncology · 2025

136 Background: The ECOG performance status (PS) is a critical measure for evaluating functional capacity and determining clinical trial eligibility in patients with advanced prostate cancer. We aim to assess the capability of large language models (LLMs) to accurately infer ECOG PS from unstructured oncology notes. Methods: This retrospective study included patients with metastatic castration-resistant prostate cancer (mCRPC) receiving Lutetium-177–PSMA-617 ( 177 Lu) at the Mayo Clinic (2022-2023). Baseline ECOG PS scores were manually curated by a trained clinician from clinical notes recorded within four weeks prior to 177 Lu initiation. Structured zero-shot prompts were then used to prompt GPT-4 to infer ECOG PS scores from unstructured oncology notes. Total dataset was randomly sampled to prompt development (20%) and held-out test set (80%). Three experimental setups were employed to assess model performance: (i) utilizing unmasked clinical notes, (ii) clinical notes with only numerical ECOG PS scores masked while retaining narrative descriptive/assessment by oncologist of ECOG score, and (iii) clinical notes with both numerical scores and narrative ECOG descriptors masked. Final performance was assessed using evaluation metrics (precision, recall, F1 score) on the held-out test set. Results: A total of 240 patients (n: 50 prompt development; n: 190 test set) were included in the analysis. The median age was 70 (IQR: 66-76); a majority of patients were White (n: 229; 95%) and non-Hispanic (n: 233; 97%). The most prevalent ECOG PS score was 0 (n: 140; 58%), followed by 1 (n: 88; 37%), and 2 (n: 12; 5%). The performance was numerically higher when using unmasked notes (weighted average precision: 0.89; recall: 0.89; F1: 0.89), compared to using notes with masked scores only (weighted average precision: 0.76; recall: 0.79; F1: 0.75) and notes with both scores and narrative masked (weighted average precision: 0.57; recall: 0.53; F1: 0.48). Metrics by each ECOG PS score is outlined (Table). Conclusions: Large language models demonstrate significant potential in precisely extracting ECOG performance status from unstructured oncology notes, offering valuable support for treatment decision-making and clinical trial enrollment. However, model performance declines as reliance on implicit reporting within clinical notes increases. Future efforts focusing on prediction of ECOG using a series of notes over time will be the next step. ECOG Unmasked oncology notes ECOG Score masked only ECOG Score + ECOG narrative marked Precision Recall F1 Score Precision Recall F1 Score Precision Recall F1 Score 0 0.89 0.92 0.91 0.75 0.93 0.83 0.59 0.74 0.66 1 0.88 0.86 0.87 0.76 0.74 0.75 0.19 0.47 0.27 2 1.00 0.82 0.90 0.78 0.33 0.47 0.78 0.16 0.26

Read the paper · More papers on PaperTik