HCMANet: Hierarchical Cross-Modality Attention Network for Underwater Salient Object Detection

Yu Wang, Wenjie Li, Haowei Wen, Zhiyang Yu, Yi Xue, Hua Li · 2024

Combining RGB images with depth maps for salient object detection (SOD) has gradually become a familiar approach. However, underwater imagery suffers severe quality degradation because of challenges like scattering, light absorption, and marine snow. These challenges invalidate the applicability of conventional salient detection methods designed for natural images. To address this challenge, we propose the Hierarchical Cross-Modality Attention Network (HCMANet) for underwater salient object detection (USOD), which effectively fuses features from RGB and depth modalities to suppress underwater noise interference and generate saliency prediction maps. Specifically, the HCMANet network employs a dual-stream parallel encoder to extract features from RGB and depth maps separately. Drawing on attention mechanisms, we further propose the Multimodal Knowledge Interaction Module (MKIM) to establish long-range dependencies between RGB and depth features and generate more discriminative information, thereby effectively tackling underwater noise and fostering feature alignment. Then, we design the Fusion Decoder (FD) that decodes the aligned features step by step to obtain accurate prediction maps. Extensive comparisons with recently proposed methods on two public USOD datasets demonstrate the superiority of our methods.

Read the paper · More papers on PaperTik