Diagnostic framework to validate clinical machine learning models locally on temporally stamped data

Maximilian Schuessler, Scott L. Fleming, Shannon Meyer, Tina Seto, Tina Hernandez‐Boussard · Communications Medicine · 2025

Real-world medical environments such as oncology are highly dynamic due to rapid changes in medical practice, technologies, and patient characteristics. This variability, if not addressed, can result in data shifts with potentially poor model performance. Presently, there are few easy-to-implement, model-agnostic diagnostic frameworks to vet machine learning models for future applicability and temporal consistency. We extracted clinical data from EHR for a cohort of over 24,000 patients who received antineoplastic therapy within a distinct year. The label of this study are acute care utilization (ACU) events, i.e., emergency department visits and hospitalizations, within 180 days of treatment initiation. Our cross-sectional data spans treatment initiation points from 2010–2022. We implemented three models within our validation framework: Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). Here, we introduce a model-agnostic diagnostic framework to validate clinical machine learning models on time-stamped data, consisting of four stages. First, the framework evaluates performance by partitioning data from multiple years into training and validation cohorts. Second, it characterizes the temporal evolution of patient outcomes and characteristics. Third, model longevity and trade-offs between data quantity and recency are explored. Finally, feature importance and data valuation algorithms are applied for feature reduction and data quality assessment. When applied to predicting ACU in cancer patients, the framework highlights fluctuations in features, labels, and data values over time. The work in this study emphasizes the importance of data timeliness and relevance. The results on ACU in cancer patients show moderate signs of drift and corroborate the relevance of temporal considerations when validating machine learning models for deployment at the point of care. With the growing use of routinely collected clinical data, computational models are increasingly used to predict patient outcomes with the aim to improve care. However, changes in medical practices, technologies, and patient characteristics can lead to variability in the clinical data that is collected, reducing the accuracy of the results obtained when applying the computational model. We developed a framework to systematically evaluate clinical machine learning models over time that assesses how clinical data and computational model performance evolve, ensuring safety and reliability. We used our model on people with cancer undergoing chemotherapy and were able to predict emergency department visits and hospitalizations. Implementing frameworks such as ours should enable the accuracy of computational models to be assessed over time, maintaining their ability to predict outcomes and improve care for patients. Schuessler at al. present a diagnostic framework for the local and temporal validation of clinical machine learning models and use to predict acute care utilization in cancer patients. It encompasses model performance evaluation, temporal evolution of features and outcomes, training schedules, and strategies for data reduction and valuation.

Read the paper · More papers on PaperTik