Vision Transformers for Medical Image Segmentation Using Breast Ultrasound Images

Gvs Raj Kumar, B. Raja Koti, Venna Mythilee Priyanka, Neeta Kumari, Eali Stephen Neal Joshua · 2025

ViT (Vision Transformer) has emerged in recent years as a potential substitute for traditional architectures of convolution, especially with more complex applications of biomedical image segmentation tasks, like tumor detection and organ segmentation. Ultrasound (US) imaging has grown in importance as a means of detecting breast cancer due to its portability, low cost, and lack of risk. Convolutional neural networks (CNNs) have dominated vision task implementations last decade. Also, they’ve shown great promise in automatically categorizing US images. Despite their achievements, they lack a comprehensive understanding of the global environment due to the limited scope of their local open region. The notion that various visual system components can maintain several foci simultaneously is the basis of Vision Transformer (ViT) designs. As an alternative to CNNs, they have recently shown great potential. Using ViT, we categorize breast ultrasound images showcasing various breast implant kinds in this study. Classification success is represented as a percentage compared to the most sophisticated CNNs, and Area Under the Curve (AUC) measures are employed to quantify performance. According to the findings, ViT models perform as well as, or even better than, CNN models when categorizing US breast images.

Read the paper · More papers on PaperTik