CGCMFNet: Context-Guided Cross-Modal Fusion Network for RGB-T Salient Object Detection
Songling He, Bin Feng Yang · 2024
Thermal infrared (T) images which are insensitive to the changes of thermal are introduced into salient object detection (SOD), termed RGB-T SOD, to overcome the performance degradation of RGB SOD in low-light or dark scenes. However, existing RGB-T SOD methods mostly integrate multimodal information by designing artificial fusion strategies, overlooking that modality differences may lead to suboptimal fusion effects. To address this issue, we propose a novel context-guided cross-modal fusion network (CGCMFNet) for RGB-T SOD, utilizing context information to reduce modality differences for intrinsic consistency in feature fusion. Firstly, we introduce a multi-scale attention module (MAM) to localize salient regions and furnishing comprehensive global context support. Subsequently, we devise a context- Guided module (CGM) that utilizes global context information as the primary guidance to mitigate modality discrepancies between two modalities. Then, a cross-modal fusion (CMF) module is proposed to merge cross-modal features exhibiting slight modality differences. Finally, through a multi-level aggregation (MA) module, fused multi-level cross-modal features are aggregated to generate high-quality salience maps. Experiments conducted on three datasets illustrate that the proposed approach yields satisfactory outcomes when compared to other state-of-the-art SOD methods.