Synergistic Fusion Network of Microscopic Hyperspectral and RGB Images for Multi-Perspective Segmentation

Lixin Zhang, Qian Wang · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Accurate segmentation of diverse structures in pathological images is crucial for medical analysis. While widely used RGB images offer high spatial resolution, microscopic hyperspectral images (MHSIs) provide unique biomedical spectral signatures. Existing multi-modal segmentation methods, however, often suffer from insufficient uni-modal learning, ineffective cross-modal interaction, and nonadaptive multi-modal fusion. Therefore, we propose a novel synergistic multi-modal learning paradigm for co-registered RGB-MHSIs, instantiated within the Synergistic Fusion Network (SyFusNet) which comprises: modality-specific modules and objectives to ensure uni-modal feature extraction, the Mutual Knowledge Sharing Module (MKSM) for explicit cross-modal interaction, and the Adaptive Dual-level Co-decision Module (ADCM) for collaborative multi-modal segmentation. Alongside uni-modal learning, MKSM disentangles MHSI- and RGB-specific features into band- and position-aware guidance, respectively, sharing as cross-modal knowledge to enhance each other’s representations. To fuse multi-modal predictions, ADCM generates global attention from integrated multi-modal features to adaptively refine decision-level outputs, yielding reliable segmentation. Experiments demonstrate that SyFusNet outperforms state-of-the-art methods with statistical significance (p< 0.01), achieving relative IoU gains of 9.35%, 4.63%, and 2.47% on the public PLGC, MDC, and WBC datasets, respectively, while also exhibiting strong generalizability and diagnostic potential through practical applications in multi-class segmentation and tumor regression grading.

Read the paper · More papers on PaperTik