Clinically Oriented Evaluation of Transfer Learning Strategies for Cross-Site Breast Cancer Histopathology Classification
Liana Stanescu, Cosmin Stoica-Spahiu · Applied Sciences · 2025
Background/Objectives: Breast cancer diagnosis based on histopathological examination remains the most reliable and widely accepted approach in clinical practice, despite being time-consuming and prone to inter-observer variability. While deep learning methods have achieved high accuracy in medical image classification, their cross-site generalization remains limited due to differences in staining protocols and image acquisition. This study aims to evaluate and compare three clinically relevant adaptation strategies to improve model robustness under domain shift. Methods: The ResNet50V2 model, pretrained on ImageNet and further fine-tuned on the Kaggle Breast Histopathology Images dataset, was subsequently adapted to the BreaKHis dataset under three clinically relevant transfer strategies: (i) threshold calibration without retraining (site calibration), (ii) head-only fine-tuning (light FT), and (iii) full fine-tuning (full FT). Experiments were performed on an internal balanced dataset and on the public BreaKHis dataset using strict patient-level splitting to avoid data leakage. Evaluation metrics included accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, computed per magnification level (40×, 100×, 200×, 400×). Results: Full fine-tuning consistently yielded the highest performance across all magnifications, reaching up to 0.983 ROC-AUC and 0.980 sensitivity at 400×. At 40× and 100×, the model correctly identified over 90% of malignant cases, with ROC-AUC values of 0.9500 and 0.9332, respectively. Head-only fine-tuning led to moderate gains (e.g., sensitivity up to 0.859 at 200×), while threshold calibration showed limited improvements (ROC-AUC ranging between 0.60–0.73). Grad-CAM analysis revealed more stable and focused attention maps after full fine-tuning, though they did not always align with diagnostically relevant regions. Conclusions: Our findings confirm that full fine-tuning is essential for robust cross-site deployment of histopathology AI systems, particularly at high magnifications. Lighter strategies such as threshold calibration or head-only fine-tuning may serve as practical alternatives in resource-constrained environments where retraining is not feasible.