Hierarchical Weighting Network with Depth Cues for Real-Time Video Salient Object Detection
Xu Mengnan, Gao Pan, Ziyan Zhang, Ping Zhang · 2022 19th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP) · 2022
Video salient object detection (VSOD) is an essential component in computer vision tasks, and depth information can provide critical spatial contour information. However, there is a lack of methods to integrate depth information into VSOD tasks. In this paper, a real-time hierarchical weighting network with depth cues is proposed to extract the features and temporal information efficiently, and depth estimation model is used to generate depth maps. In order to fuse the features of different modalities, a cross-modal weighting (CMW) module based on three-dimensional convolution is proposed, which enables the powerful aggregation capability. Furthermore, an efficient temporal information query module (TQM) with Transformer attention mechanism is designed to extract crucial temporal information from video sequences. Experiments on public VSOD datasets show that our method is excellent in detection speed and the results are competitive in the state-of-the-art methods.