Cross Modal Adaptive Feature Fusion Network for RGB-D Salient Object Detection

Haitang Li, Lijuan Shi, Yanjun Wang, Yuanfei He, Weihua Zhou, Peiling Li · Research Square · 2024

Abstract RGB-D salient object detection is a fundamental task in computer vision, garnering considerable attention in recent years as many visual tasks rely on salient object detection as a foundational step. However, in the context of RGB-D salient object detection, the effective fusion of RGB modality features with DEPTH modality features to enhance the performance of the salient object detection task remains a significant challenge. In this paper, addressing this issue, we propose an effective network architecture, CMAFF-Net: We design a novel attention mechanism called Normalized Coordinate Attention (NCA) to further extract features from the backbone network, resulting in more effective feature representation. We introduce the CMAFF module to adaptively fuse features extracted at different stages of the backbone network, ensuring the effectiveness of feature fusion at each stage. We propose a hybrid loss function to provide more effective supervision for model learning. Our model exhibits superior performance on the first four publicly available datasets compared to the latest methods. Particularly noteworthy is the approximately 4% improvement in the F-measure evaluation metric on the SSD dataset.

Read the paper · More papers on PaperTik