CCEnd-Net: Cross-Modal Cascaded Encoder-Decoder Network for Multisource Data Fusion Classification
Song Dai, Dongmei Song, Bin Wang, Weimin Chen · IEEE Transactions on Geoscience and Remote Sensing · 2025
Multi-source data fusion offers great potential for land cover classification. However, the substantial differences in data structures and content representations across various remote sensing sources present significant challenges in heterogeneous feature extraction and information fusion, ultimately constraining the effectiveness of fusion-based classification. To address the aforementioned limitations, this article proposes a multi-source data fusion classification method based on a cross-modal cascaded encoder-decoder network (CCEnd-Net). The proposed algorithm comprises two primary components: a blur feature extraction module and a cascaded encoder-decoder feature fusion module. Specifically, the blur feature extraction module utilizes multi-dimensional convolution to extract deep features from multi-source data and incorporates a blur pooling module to enhance aliasing resistance. This approach mitigates the original spatial discrepancies among heterogeneous data while preventing feature distortion. Meanwhile, the cascaded encoder-decoder feature fusion module reconstructs multi-source data features by integrating a multi-scale channel-spatial interaction attention (MCIA)-enhanced convolutional neural networks (CNNs) with a transformer-based cross-modal fusion (TCMF) block. Additionally, a multi-level fusion strategy is employed to comprehensively exploit the complementary information from different remote sensing data sources, thereby improving classification performance. To validate the effectiveness of the proposed method, comprehensive experiments were conducted on four benchmark datasets covering light detection and ranging (LiDAR), hyperspectral imaging (HSI), synthetic aperture radar (SAR), and very high resolution (VHR) data. The proposed CCEnd-Net was systematically evaluated against state-of-the-art models, including transformer-based architectures, CNNs, and traditional classifiers. The experimental results demonstrate that CCEnd-Net achieves superior performance across all datasets.