Video Salient Object Detection Network with Bidirectional Memory and Spatiotemporal Constraints
Hongyu Wang, Nan Mu, Yu Zhang · 2021 IEEE International Conference on Systems, Man, and Cybernetics (SMC) · 2021
Although deep learning-based video salient object detection networks have shown outstanding performance, there still exist problems such as the incompleteness of salient objects due to the lack of spatial saliency information, and the globally inconsistency of saliency results detected in long-term video sequences. To address these issues, we propose a novel video salient object detection network with bidirectional memory and spatiotemporal constraints. The global and local cues are first cascaded to extract spatial location information and internal structural details by residual skip connections. After obtaining reliable and fine-grained spatial saliency for a single frame, we design a DB-ConvLSTM module with bidirectional memory that retains the past and future information. A global temporal attention mechanism is added in the temporal dimension to correlate each frame with the whole video to obtain accurate salient objects of globally consistent. Extensive experiments have been conducted to demonstrate that the proposed model outperforms the other seven state-of-the-art models on four datasets, viz. DAVIS, FBMS, SegTrackV2, and UVSD.