Explicit Semantic Alignment Network for RGB-T salient object detection with Hierarchical Cross-Modal Fusion

Hongkuan Wang, Qifeng Yu, Zhenguang Di, Gang Yang · Image and Vision Computing · 2025

Existing RGB-T salient object detection methods primarily rely on the learning mechanism of neural networks to perform implicit cross-modal feature alignment, aiming to achieve complementary fusion of modal features. However, this implicit feature alignment method has two main limitations: first, it is prone to causing loss of the salient object’s structural information; second, it may lead to abnormal activation responses that are not related to the object. To address the above issues, we propose the innovative Explicit Semantic Alignment (ESA) framework and design the Explicit Semantic Alignment Network for RGB-T Salient Object Detection with Hierarchical Cross-Modal Fusion (ESANet). Specifically, we design a Saliency-Aware Refinement Module (SARM), which fuses high-level semantic features with mid-level spatial details through cross-aggregation and the dynamic integration module to achieve bidirectional interaction and adaptive fusion of cross-modal features. It also utilizes a cross-modal multi-head attention mechanism to generate fine-grained shared semantic information. Subsequently, the Cross-Modal Feature Alignment Module (CFAM) introduces a window-based attention propagation mechanism, which enforces consistency in scene understanding between RGB and thermal modalities by using shared semantics as an alignment constraint. Finally, the Semantic-Guided Edge Sharpening Module (SESM) combines shared semantics with a weight enhancement strategy to optimize the consistency of shallow cross-modal feature distributions. Experimental results demonstrate that ESANet significantly outperforms existing state-of-the-art RGB-T salient object detection methods on three public datasets, validating its excellent performance in salient object detection tasks. Our code will be released at https://github.com/whklearn/ESANet.git .

Read the paper · More papers on PaperTik