Enhanced Breast Cancer Detection Using a Hybrid Vision Transformer and CNN Model with Optimized Feature Integration
Uday Pratap Singh, Bersha Kumari, Ebtasam Ahmad Siddiqui, Adil Tanveer · 2024
Detection of breast cancer in medical image has become a challenge in aspects of precision and reliability. This research proposes a new hybrid model composed of Vision Transformer (ViT) and Convolutional Neural Network (CNN) architectures, enhancing the effectiveness of local and global feature extraction. Moreover, the model combines CT and MRI modalities through multi-feature fusion to further improve performance. The model has optimal hyperparameters, such as a learning rate of 0.002, batch size of 40, and transformer depth of 12, which were determined by, ablation studies, and sensitivity analysis. The proposed hybrid model outperforms other conventional methods with a 99.1% accuracy. Our findings prove the efficiency of hybrid architectures in breast cancer diagnosis with improved specificity and sensitivity.