Learning in Medical Image Segmentation: SAM Fine-Tuning vs. U-Net Training for Breast Ultrasound

Laila El Jiani, Oussama Kaich, Zakaria El Fakir, Sanaa El Filali, El Habib Benlahmar, Fatima Zahra-Alaoui, Zouheir Banou, Amine Mounjid · 2025

Breast cancer diagnosis and treatment planning rely on precise tumor segmentation in medical imaging. Deep learning models, particularly CNNs and Transformer-based architectures, have shown strong potential in automating this task. This study compares U-Net, a widely used CNN-based model, with the Segment Anything Model (SAM), a Vision Transformer pre-trained on natural images, for breast ultrasound tumor segmentation. While U-Net was trained from scratch on a breast ultrasound dataset, SAM was fine-tuned using prompt-based learning with bounding box annotations. The models were evaluated using IoU, Dice Similarity Coefficient (DSC), accuracy, precision, and recall. Experimental results show that fine-tuned SAM achieved an IoU of 0.9311, surpassing U-Net while maintaining competitive computational efficiency. Additionally, SAM’s ability to generalize across different breast tumor types highlights the potential of transformer-based architectures in medical imaging.This study provides insights into the trade-offs between CNN-based and transformer-based models for ultrasound tumor segmentation, emphasizing the impact of transfer learning and fine-tuning in adapting large-scale models to specialized medical applications.

Read the paper · More papers on PaperTik