Y-Net: A Compact Edge-Friendly Model for Multi-Modal Medical Image Segmentation
Soma Dasgupta, Swarnava Dey, Avik Ghose, Arijit Mukherjee, Arpan Pal · 2025
Multi-modal image segmentation has significant potential for advancing high-quality representation learning, as different modalities provide complementary information about anatomical structures, organs, and diseases. However, there is currently no principled approach to designing compact, edge-efficient architectures that effectively leverage multi-modal images for medical image analysis. Existing practices are either manual, relying on expert-driven fusion of features from unregistered and unpaired modalities, or employ overly large architectures with multiple encoders and decoders.In this paper, we bridge this gap by introducing Y-Net, a novel architecture tailored to jointly learn segmentation tasks for multiple organs using data from diverse medical imaging modalities. Y-Net, combined with an automated hyperparameter search methodology, can be deployed in a plug-and-play fashion on multi-modal imaging datasets, delivering accurate and parameter-efficient organ segmentation. We validate Y-Net on the CHAOS challenge task, which involves segmenting abdominal organs from CT and MRI data. Our approach achieves a 6% improvement in Intersection over Union (IoU) scores across all organ classes compared to state-of-the-art single-modality segmentation methods, while requiring only one-twentieth of the parameters. This compactness makes Y-Net particularly well-suited for on-premises, privacy-preserving inference in healthcare analytics.