CMDNet: salient object detection for RGB-D video based on cross-modal feature transmission
Jianlin Guo, Zhiqiang Lu, Luwang Li · 2025
Depth information provides essential geometric cues for Salient Object Detection (SOD) and plays a key role in RGB-D methods. However, Video Salient Object Detection (VSOD) has largely overlooked depth, focusing instead on spatiotemporal features. To address this gap, we introduce CMDNet, a multimodal network that integrates RGB and depth information. CMDNet employs two novel modules: the Cross-Modality Feature Adapter (CMFA) for dynamic feature transfer and the Context-Aware Attention Module (CAAM) for efficient feature fusion. These components enhance accuracy and robustness, as validated by experiments on five benchmark datasets, where CMDNet outperforms state-ofthe-art methods.