HybridMammoNet: A Hybrid CNN-ViT Architecture for Multi-view Mammography Image Classification

Hind Allaoui, Youssef Alj, Yassine Ameskine · 2024

In recent years, Computer-Aided Diagnosis (CAD) from mammography images has raised the interest of numerous researchers in the deep learning field. However, the majority of existing methods and architectures either rely on a single view mammography approach or employ late fusion methods to combine CNN-features from both views. Furthermore, most of these methods predominantly employ pure architectures, such as CNN-based or Transformers-based models, which limit their ability to compute both local (via CNNs) and global (via Transformers) features simultaneously. To address these limitations, we introduce an architecture designed to accept input from two views and merge them using an attention layer beyond the feature map level, where local features are computed. To validate our approach, we use the CBIS-DDSM public dataset, which includes two views (CC and MLO) for each breast of every patient. Our results suggest that employing a hybrid model capable of calculating various types of features from both views before merging them outperforms the utilization of a late-joining CNN-based architecture. This approach represents a promising solution for enhancing mammographic image classification by effectively combining local and global features from multiple views.

Read the paper · More papers on PaperTik