Self-Supervised Pretraining With Multimodality Representation Enhancement for Salient Object Detection in RGB-D Images

Lina Gao, Bing Liu, Ping Fu, Mingzhu Xu, Yonggang Zhang, Yulong Huang · IEEE Transactions on Instrumentation and Measurement · 2025

Training self-supervised pretraining models for salient object detection (SOD) in RGB-D images is appealing, as it removes the costly demand of explicitly pixel-wise labels and exhibits promising saliency performance. However, the limited discriminability of multimodality representation indicates significant potential for developing a multimodality representation enhancement pretraining model. In this article, we focus on designing a self-supervised pretraining model to learn multimodality representations from unannotated data, which consists of an appearance representation learning (ARL), a depth enhancement self-supervised learning (DESL), and a contrastive multimodality representation learning (CMRL) auxiliary task. The former two pretext tasks can learn additional representation cues and spatial structure information for saliency detection, while the latter is designed to further enhance inter-modal compactness and correlation. Furthermore, we propose a holistic RGB-D SOD model with a cross-modality information fusion (CMIF) module to learn inter-modal complementary features and a hierarchical interaction aggregation module (HIAM) to capture the global context relationship at inter-level. Extensive experiments have demonstrated the effectiveness of the proposed pretrained model on the RGB-D SOD task. The fine-tuned model can achieve the highest self-supervised performance currently available and comparable results with state-of-the-art fully supervised approaches on eight RGB-D SOD testing datasets. Additionally, the proposed model has been found to reduce the annotation effort by at least 53%, making it a cost-effective alternative to fully supervised models.

Read the paper · More papers on PaperTik