AI-Driven Feature-Enhanced Stacking Ensemble With Global-Context Vision Transformers for Breast Cancer Classification in Ultrasound Images
Nghia Trong Vo, Hoang Phi Yen Duong, Tuan Thanh Nguyen, Nhan Duc Le, Trung Q. Duong · IEEE Internet of Things Journal · 2026
Breast cancer remains a leading cause of death among women worldwide. Early detection of breast cancer is a crucial step towards improving survival rates for patients affected by the disease and is typically performed with the help of ultrasound imaging. Current rapid advancements in artificial intelligence (AI) research have produced a plethora of machine learning methods that aid in building automated diagnostic assistance systems for early cancer detection, including breast cancer detection. While deep learning has shown promise in medical image analysis, most existing approaches rely on single models or simple ensemble methods that fail to fully exploit complementary feature representations across architectures. This paper introduces a novel feature-enhanced stacking ensemble framework that combines state-of-the-art global context vision transformer (GCViT) with well-established convolutional neural network (CNN) architectures (ResNet-50V2, ConvNeXt-Tiny, and EfficientNetV2-B3) for automated breast cancer classification from ultrasound images. Unlike conventional ensembles that aggregate only prediction probabilities, our approach extracts deep feature embeddings from a dedicated CNN branch and concatenates them with base model predictions as input to a meta-learner, a multi-layer perceptron (MLP), enabling the ensemble to leverage both decision-level and feature-level information. When incorporating a meta model with feature representations from a CNN-based feature extractor, we are able to produce superior performance across multiple metrics compared to prior works. We accomplish top performance of 94.23% accuracy, 95.47% AUC-ROC. To further evaluate the robustness and generalizability of our approach, we conduct additional experiments on the melanoma cancer image dataset and achieve 95.4% accuracy. We provide comprehensive explainability analysis through shapley additive explanations (SHAP) values for feature attribution, permutation importance for model contribution quantification, and saliency maps for visual interpretation from base models and the end-to-end ensemble model to explain their contributions to final predictions.