Modal-Aware Interaction Network for RGB-D Salient Object Detection

Longsheng Wei, ZiQiang Zhu · IEEE Transactions on Instrumentation and Measurement · 2025

RGB-D salient object detection (SOD) focuses on identifying the most important regions in an image by leveraging the complementary fusion of RGB and depth features. Most existing methods adopt a dual-stream approach at the encoding stage to extract feature information using a backbone network, followed by the fusion of these multi-modal features with a common module. However, these methods often overlook the differences in the scale of features extracted by the backbone network. In this paper, we introduce a modal-aware interaction network (MINet) designed to improve the efficiency of feature fusion between RGB features from a dual-stream backbone network and depth features. Our approach aims to enhance the complementarity between RGB and depth features in RGB-D SOD tasks. Specifically, we build a modality perception space based on features of varying dimensions extracted during the feature extraction phase of the dual-stream backbone. We also propose a detailed supplementation (DS) module to refine edge features by enhancing low-level feature information, and a guided semantic awareness (GSA) module to fuse high-level modal features in a complementary manner, reducing background redundancy and effectively locating salient objects. Furthermore, in our proposed advanced semantic aggregation (ASA) module, global information is integrated into contextual features to explicitly guide the up-sampling process, aiding the precise localization of salient targets. The entire up-sampling process is supervised by a hybrid loss function. Our modality-aware interaction space successfully balances the processing of high-level spatial location features and low-level texture details, extracting more detailed details while capturing spatial structure, resulting in an improved feature fusion process. Experimental results demonstrate that our model is not only simple to use but also outperforms 19 existing methods across 7 public datasets on 4 evaluation metrics.

Read the paper · More papers on PaperTik