SwinMatcher: Universal Cross-Modal Remote Sensing Image Matching With Interactive Swin Transformer
Wei Li, Desheng Weng, Chenzhong Gao, Qian Du · IEEE Transactions on Geoscience and Remote Sensing · 2025
Cross-modal remote sensing image matching serves as a key technique for collaborative utilization of multi-source information. However, modal differences and geometric distortions between multi-source images pose challenges to existing methods in terms of robustness and generalization. To achieve feature interaction in cross-modal scenes, this paper proposes SwinMatcher, an end-to-end matching model based on the Transformer architecture. Innovatively proposing the window/shifted-window cross-attention based on the window/shifted-window self-attention mechanisms of Swin Transformer, SwinMatcher enables efficient cross-modal feature interaction and multi-scale contextual modeling. It also incorporates a learnable matching module to directly generate semi-dense correspondences. Moreover, a cross-modal remote sensing image matching dataset is generated, which encompasses four modalities: visible light, synthetic aperture radar (SAR), light detection and ranging (LiDAR), and map, distributed across four representative scenes. The dataset includes 400 samples produced via random homography transformations, designed to enhance modal diversity and scene complexity. Experiments demonstrate that SwinMatcher outperforms state-of-the-art methods on this new dataset as well as public benchmarks, exhibiting superior robustness under complex scenes involving coupled modal and geometric distortions. The proposed method and dataset provide novel solutions and evaluation benchmarks for cross-modal remote sensing image matching. The code and testing dataset will be made publicly available at https://github.com/LotrL/SwinMatcher.