State Space Model-Based Fusion Modulation Network for Multimodal Semantic Segmentation

Yanli Shang, Fanghong Liu, Haowen Zhang, Qiuze Yu · IEEE Geoscience and Remote Sensing Letters · 2025

Combining synthetic aperture radar (SAR) and optical images for geomorphic feature extraction offers a promising solution. However, the significant difference in data distribution between the two modalities poses a challenge for the fusion mechanism to achieve efficient segmentation. This letter presents a fusion modulation network based on state space model (SSM) (FMamba), designed for multimodal semantic segmentation tasks. The network can effectively fuse the complementary features of the two modalities to enhance the semantic features of the target regions. Specifically, FMamba integrates a dynamic fusion modulation module (DFMM) based on SSM at the encoding stage. This module dynamically fuses and modulates features from two different modalities to facilitate cross-modal feature extraction and semantic alignment. In the decoding phase, the final segmentation map is generated by progressively decoding and reconstructing the multiscale fusion features. Experimental results on the WHU-OPT-SAR and DFC2025 datasets demonstrate the superior semantic segmentation performance of the proposed method.

Read the paper · More papers on PaperTik