A multimodal remote sensing image fusion method for object detection based on tensor decomposition

Yunqi Zhao, Ke Liu, Dongdong Lu, Zhe Zhang · 2025

As one of the most important and challenging tasks in the field of computer vision, object detection has long played a huge role in aerospace, autonomous driving, industrial inspection and other fields. Traditional object detection usually uses optical (RGB) images as the original data, but with the progress of sensor technology, the modalities of available data have increased. Especially in the aerospace field, the types of remote sensing images are becoming diverse, including infrared (IR) images, SAR images, hyperspectral images and so on. Compared with single-modal, multi-modal remote sensing images can reflect different attributes of the same observed object or scene, and these attributes are usually complementary. Many existing studies have indicated that multi-modal remote sensing images are significantly better than single-modal images in many tasks. The multi-modal image fusion methods are mainly divided into two categories: fusion based on deep learning and fusion based on tensor analysis. Deep learning based fusion mainly includes three types: early fusion, mid-term fusion, and decision-level fusion. In recent years, fusion methods based on deep learning have achieved certain success in object detection tasks. Fusion based on tensor analysis regards multi-modal images as high-order tensors and uses tensor decomposition, such as Tucker decomposition and CP decomposition, to obtain interaction information between modalities. Some studies have shown the potential of tensor based methods in multi-modal data fusion and achieved good performance in tasks such as emotion recognition and landuse classification. However, there is still limited research on applying tensor based fusion methods to multi-modal object detection tasks. In this paper, we propose a remote sensing image fusion method based on Tucker decomposition and apply this method to early fusion and mid-term fusion of deep neural networks. The effectiveness of this method has been validated in an excellent remote sensing image object detection deep learning model. After adding tensor methods, the detection precision of early fusion increased from 65% to 75%, and the detection precision of mid-term fusion increased from 76% to 79% and 82%.

Read the paper · More papers on PaperTik