Salient Object Detection Based on High-level Semantic Guidance and Multi-modal Interaction
Chao Yang, Zheng Guan, Wenbi Ma · 2024
Semantic information is essential in RGB-T salient object detection (SOD). Most existing methods directly input the extracted low-level features into the interaction module and utilize a simple recursive structure for prediction. Despite their excellent performance in several scenarios, they suffer from capturing and exploiting the attributes and complementary potential between different feature layers of images, which are critical in obtaining details and object location. In this work, we proposed a top-down SOD method based on high-level semantic guidance information and multi-modal interaction information. On the one hand, a high-level semantic module (HSM) is designed to implement high-level semantic guidance to the decoder part and to retain the location information of salient objects. On the other hand, a multi-modal interactive module (MIM), which composed of channel attention and multi-branch structure, is designed to interact cross-modal features and retain object location information. Evaluation results on three common benchmark datasets reveal that the proposed method successfully achieves competitive state-of-the-art performance.