Implementation of Explainable AI on CNN- and ViT-Based Models for Classifying Breast Cancer
Antoluis Hendry, Erika Felicia, Yonatan Natanael, Reinert Yosua Rumagit · Procedia Computer Science · 2025
The use of deep learning in the medical area makes many contributions, especially in aiding the medical staff to diagnose and identify global disease. However, those models still lack transparency in their decision-making process, which affects the trust between the medical staff and the deep learning models. This paper aims to compare the performance of ViTs (ViT/B-16, Swin-S, and DeiT-B) and CNNs (ResNet50, DenseNet121, and EfficientNet-B0)-based models and the result of XAI techniques (SHAP, LIME, Attention rollout, and Grad-CAM) implementation for better interpretability of deep learning models in breast cancer ultrasound dataset classification. The result showed competitive performance between ViT and CNN. Based on the model evaluation result, Deit has the most balanced performance for classifying all 3 classes (normal, benign, and malignant), with an average of 0.95 for precision, recall, and F-1 score for the normal class; 0.88 for the benign class; and 0.78 for the malignant class. On the other hand, EfficientNet scores 0.85 recall and 0.79 F1-score for the malignant class, implying EfficientNet has the best performance at classifying the malignant class amongst the whole model. The implementation of XAI techniques also has its strengths and weaknesses, and a hybrid approach is suggested for better clarification. With attention rollout or Grad-CAM, quick localization can be accomplished. LIME or SHAP can be used to confirm and enhance these results with rigorous model-agnostic attribution. SHAP works best when precise, class contribution scores are required for any model architecture, while attention rollout is the ideal method for Transformer-based models that favor boundary precision.