Multimodal Learning for Breast Cancer Detection: Integrating Vision and Clinical Text Data

Maryam Alaei, Mohammad Zare, M. Hazrati, Amir Chekini · International Journal of Science and Research (IJSR) · 2025

Early and accurate detection of breast cancer is critical for improving patient outcomes and reducing mortality. In this paper, we propose a multimodal deep learning framework that integrates high-resolution mammographic image analysis using Vision Transformers (ViTs) with clinical text interpretation through BERT-based language models. By combining visual and textual information via an early fusion strategy, our approach captures complementary diagnostic cues to enhance prediction accuracy. We evaluate our model on two publicly available datasets-CBISDDSM and MIMIC-CXR-and demonstrate that the multimodal system significantly outperforms unimodal baselines. Our best-performing model achieves an accuracy of 91.4 and an AUROC of 0.94, surpassing both ViT-only and BERT-only models. Additional experiments and ablation studies confirm the effectiveness of the fusion strategy and the contribution of each modality. These findings highlight the potential of multimodal transformer-based learning to support radiologists in early breast cancer diagnosis through more holistic and robust decision-making.

Read the paper · More papers on PaperTik