MoSA: Modality-Aware Spatially-Adaptive Adaptation for RGB-X Semantic Segmentation

Xing Zhang, Kunlei Dong, Wei Wang, Huasheng Yang, Lei Chen, Zhengpin Li, Jian Zhang · IEEE Access · 2026

Multimodal semantic segmentation enhances semantic perception by fusing complementary features from multiple modalities. However, the contribution of each modality to segmentation predictions varies across spatial locations depending on local content characteristics, an issue that has been insufficiently addressed by existing methods. To tackle this challenge, we present MoSA, a novel framework for adapting vision foundation models to Multimodal semantic segmentation via Modality-aware Spatially-Adaptive learning. MoSA introduces a Modality-Specific Spatially-Modulated Adapter (MS-SMA) to capture modality-specific representations while dynamically adjusting adaptation intensity based on both global scene context and local content information. Furthermore, MoSA proposes Reliability-Guided Cross-Modal Fusion (RGCF), which estimates per-location modality reliability using feature statistics and performs adaptively weighted fusion. MoSA is a model-agnostic framework that can be seamlessly applied to various multimodal combinations without task-specific modifications. MoSA outperforms state-of-the-art methods in mIoU by 2.3% on the NYU Depth V2 dataset and by 2.1% on the MFNet dataset, requiring only 7.7% of the trainable parameters.

Read the paper · More papers on PaperTik