Use of an artificial intelligence model to predict Ki67 from H&E-stained whole slides images in breast cancer.
Daniel Kates-Harbeck, Martin Filipits, Hans Heinrich Kreipe, Dominik Hlauschek, Matthias Christgen, Gabriel Rinnerthaler, Oleg Gluz, Simon Peter Gampenrieder, Sven Mahner, Karin Haider, Wolfgang Hulla, Ronald Ernest Kates, Jingbin Zhang, Alexander Piehler, Hans Pinckaers, Gijs Smit, Jacqueline Reedy Griffin, Nadia Harbeck, Michael Gnant · Journal of Clinical Oncology · 2025
e13652 Background: Accurate assessment of Ki67 is critical for evaluating cellular proliferation and tumor aggressiveness in breast cancer diagnosis and prognosis. Traditionally, Ki67 immunohistochemistry (IHC) requires appropriate pre-analytical handling, standardized visual scoring, and experienced pathologists for correct interpretation. IHC is a laboratory-intensive procedure that can be affected by inter-observer variability (IOV) and residual heterogeneity. Methods: In this study, we introduce a novel artificial intelligence (AI) model that predicts Ki67 directly from Hematoxylin and Eosin (H&E)-stained Whole Slide Images (WSIs) in patients with hormone receptor-positive early-stage breast cancer. Our model utilizes deep-learning techniques to identify histopathological features that correlate with Ki67 in the whole tissue sample. This makes it hotspot-independent and enables accurate Ki67 predictions for heterogeneous tumor regions and a more comprehensive assessment. The AI-model was developed and validated using over 5200 patients from WSG ADAPT HR+/HER2- and PlanB trials with a 60/40 split and externally validated in ABCSG 6 (N = 1115). The primary objective of the Ki67 model is to correctly classify whether a tumor has low Ki67 (< 20%) or high Ki67 (≥20%) based on pathologist-annotated ground truth, defined as baseline Ki67 assessment by IHC. The area-under-the-curve (AUC) and associated 95%-confidence intervals (CI) were used to evaluate discrimination. To illustrate clinical utility, an exploratory decision threshold was estimated by maximizing the Youden Index in each validation dataset, providing estimates of sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Results: The AI-powered Ki67 classifier model demonstrated an AUC of 0.811 (95% CI: 0.791-0.827) in the test split and 0.842 (95% CI: 0.818-0.868) in external validation. Using the exploratory optimal threshold of 0.418 in the test split, the sensitivity and specificity for determining high versus low Ki-67 expression were 77% and 70% respectively. Using the exploratory optimal threshold of 0.616 in the external validation dataset, sensitivity was 73% and specificity was 81%. In the test split and external validation dataset, the PPV was 72% and 64% respectively, and NPV was 75% and 87% respectively. Conclusions: This novel approach to classifying Ki67 directly from H&E-stained slides offers a promising automated solution to streamline diagnostic workflows and enables more accurate, reproducible Ki67 assessment. Ongoing efforts are focused on refining and validating this model to enhance its clinical utility and applicability.