A Semi-Supervised Image Registration Framework Based on Multimodal Cross-Attention

Ming Zhao, Jingyi Liu, Yan Wu · IEEE Geoscience and Remote Sensing Letters · 2024

Registration of multimodal image pairs is a fundamental task in many remote sensing applications. In order to achieve accurate and low-cost remote sensing image registration, we propose a semi-supervised image registration framework based on multimodal cross-attention, which consists of the encoder for feature extraction, multimodal cross-attention module, and detection/descriptor decoders. We adopt positional encoding for feature maps to enhance the features with spatial contexts, especially for remote sensing images with large geometrical deformations. In order to learn common features that independent of modalities between multimodal images, we proposed multimodal cross-attention module to extract cross modal features, which helps the detectors to extract more reliable matching keypoints. The network is trained in a semi-supervised manner, which requires only a small dataset of incompletely labeled images. In order to learn reliable keypoints from image pairs with inconsistent intensity and geomitrical deformations, we randomly establish different geometrical mappings for the multimodal image pairs during training, and then enrich the keypoint labels by continuously adding reliable keypoints extracted by the detection decoder in each epoch. Experimental results show that the proposed method achieves more comprehensive and accurate registration than the state-of-the-art methods for multimodal remote sensing images.

Read the paper · More papers on PaperTik