SegMAE-Net: A Hybrid Method Using Masked Autoencoders for Consistent 3D Medical Image Segmentation
Zheng Kai Liaw, Ankit Das, Shaista Hussain, Feng Yang, Yong Liu, Rick Siow Mong Goh · 2024
Volumetric medical image segmentation has been particularly challenging as both local and global features are important in producing an accurate and consistent segmentation output. However, 2D CNNs often ignore the global contextual information of the volumetric input while the use of 3D CNNs is heavily limited by the large computational needs and GPU restrictions. In this paper, we propose SegMAE-Net, which combines 2D and 3D methods to leverage on the strengths of both approaches. Specifically, SegMAE-Net consists of two branches, 1) slice-centric branch made up of an encoder-decoder architecture to learn the local features of each image slice, 2) volumecentric branch utilising masked autoencoders to capture longrange dependencies. We evaluate SegMAE-Net on the RETOUCH dataset, and the experimental results show that our proposed method achieves state-of-the-art performance. We also show that our method is able to produce segmentation outputs with a higher consistency across the volume level.