Large Model Guided Semantic Segment Anything for Underwater Consumer Electronics
Qirui Lin, Hua Li, Yuheng Jia · IEEE Transactions on Consumer Electronics · 2025
With the rapid advancement of artificial intelligence (AI), large model techniques are increasingly being integrated into underwater consumer electronics, enhancing the functionality and user experience of devices used in aquatic environments. Large-scale models like Segment Anything Model (SAM) have revolutionized AI research, particularly in the field of computer vision. However, despite the powerful image segmentation capabilities of SAM, it still faces some difficulties like over-segmentation and blurry segmentation boundaries in underwater scenarios, which may affect the development of underwater consumer electronics. In this paper, we propose a generalist-specialist-like framework for underwater image semantic segmentation called Underwater Semantic Segment Anything (UWSSA). We adopt the naive SAM as a Boundary-aware Module to generate high-quality masks without semantic information. Then, we propose an efficient semantic segmentation method called Adapted Underwater Vision Transformer (AUViT), which serves as a specialist to generate predicted semantic maps with pixel-level annotations. To select accurate category labels for predicted semantic maps from masks, we introduce a SAM-ViT Interaction Module (SVIM) to filter high-quality masks and match the correct semantic information. Extensive experimental results demonstrate that our method has better performance compared to current mainstream semantic segmentation methods and SAM, achieving significant gains of 5.7% mIoU on SUIM over ViT-Adapter. We hope our method can enhance the development of underwater consumer electronics, leading to better consumer experiences and new applications.