MDFformer: A Multiscale Dual-Modal Fusion Transformer for Semantic Segmentation
Leixiong Shi, Yifei Dong, Tianqi Wang, Rui Cao, Xin Wen · IEEE Geoscience and Remote Sensing Letters · 2025
Remote sensing image interpretation is a crucial process for Earth observation and geoscience research, and multi-modal fusion techniques are key to improving semantic segmentation accuracy. CNN-based methods, due to the limited receptive field, suffer from inadequate long-range contextual modeling. While Transformer-based architectures alleviate the limitation, existing multi-modal fusion Transformers still face challenges in efficiently integrating multi-scale information. Therefore, this work proposes a multi-scale dual-modal fusion Transformer, namely MDFformer, to integrate RGB and DSM cross-modal interaction information. More specifically, MDFformer is a dual-stream encoder architecture combined with a multi-scale fusion FPN decoder. MDFformer introduce two novel modules: CSFM (Cross-Modality Shuffle Fusion Module) and SDEM (Spatial Detail Enhance Module), along with deep supervision(DS) to alter the gradient flow and enhance multi-scale feature learning. As a result, MDFformer efficiently utilizes both shallow and deep features from diverse modalities across multiple scales. The validation is performed on two publicly available high-resolution datasets, Vaihingen and Potsdam, achieving an overall accuracy (OA) of 92.06% and 91.25%, respectively. Therefore, the model effectively handles RGB and DSM data, making it suitable for remote sensing tasks and efficient earth sensing.