CMTDet: Complementary Mask-Enhanced Transformer for Multimodal Remote Sensing Object Detection

Jianfeng Li, Siyu Cheng, Xiang Li, Longqing Tu, Yingjie Mei, Guangjiao Zhou, Chenxu Wang · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · 2026

Object detection based on multi-modal images is an emerging method in remote sensing which has gained extensive attention in recent years, particularly for all-weather robust detection applications. However, in multi-modal remote sensing object detection, there are issues of over-reliance on a single modality and inter-modal interference. To solve these problems, this paper proposes a Complementary Mask-enhanced Transformer Detection network (CMTDet). Two key modules are designed in CMTDet: the Dual-stage Cross-modal Interaction (DCMI) module and the Common-mode Guided Weighted Fusion (CMGF) module. The Dual-stage Cross-modal Interaction (DCMI) module consists of two components: the Differential Guided Complementary Enhancement (DGCE) sub-module and the Complementary Mask Cross-Attention (CMCA) mechanism. The DGCE component captures differential information from features of different modalities, enhancing the ability to express the semantic-channel correlation and differential information between modalities. Then, the CMCA component explores the correspondence between features of different modalities through cross-attention and reduces the dependence of feature extraction on a single modality using asymmetric complementary masks. After the improved feature extraction, the CMGF module generates guidance signals based on common-mode information, explores common information between modalities and the contribution of information across different modalities, and designs adaptive weights to retain the advantageous information of each modality. Ultimately, this enhances the semantic expression ability and robustness of the fused modality. Extensive experiments across DroneVehicle, VEDAI, and OGSOD datasets show superior performance and generalization capability. Extensive experiments on the DroneVehicle, VEDAI and OGSOD datasets show that CMTDet outperforms the baseline: it achieves 1.87% and 1.73% higher mAP50 and mAP50-95 on the DroneVehicle dataset, with a 1.98% mAP boost in nighttime scenarios, verifying its superior detection performance.

Read the paper · More papers on PaperTik