A Multi-Modal and Multi-View Fusion Network for BI-RADS Five-Classification of Breast Tumors
Huiru Ming, Lei Yang, Ming‐Hui Chen, Suya Han, Hongwei Xu, Ling Ma, Xin Zhao, Huiqin Jiang · 2024
The precise classification of breast tumors in mammography examinations is crucial for early diagnosis and the formulation of treatment plans. The five-classification of breast tumors in Breast Imaging Reporting and Data System (BI-RADS) is a challenging task that demands a comprehensive analysis and precise classification. However, most current studies fail to meet the clinical demand of combining multi-view images and clinical textual information for the BI-RADS five-classification of breast tumors. To address these challenges and better meet clinical requirements, a fusion network combining multi-modal (text and image) and multi-view (craniocaudal and mediolateral oblique views) is proposed to realize the BI-RADS five-classification in this paper. The proposed network employs a dual-branch structure to extract image and non-image features. The first branch consists of a Transformer-based BiFormer backbone network for image features extraction, while the second branch utilizes an MLP-mixer backbone network to extract clinical text features. Additionally, we introduce a "cross modal attention block" and a "cross view attention block" for fusing multi-modal and multi-view information respectively, along with the utilization of "classification token" to gather all useful information for the final prediction. Experiments conducted on the DDSM dataset demonstrates the effectiveness of multi-modal and multi-view fusion, and achieves 87.7% accuracy.