Multi-Scale Attention and Encoder-Decoder Network for Video Saliency Object Detection
Hongbo Bi, Huihui Zhu, Lina Yang, Ranwan Wu · Pattern Recognition and Image Analysis · 2022
Abstract— In recent years, video saliency object detection has received more and more attention, and many excellent algorithms have been proposed. In the paper, we propose a new idea of video saliency object detection, named MAED-Net. Our method is mainly divided into two modules: spatial module and temporal module. In spatial module: we use a set of parallel dilated convolutions, and add channel attention to each dilated convolutions. Multi-scale mimics the characteristics of the human retina. Attention is to imitate the human attention mechanism. We combine multi-scale information with attention information, which constitutes the pyramid multi-scale channel attention. Multi-scale channel attention allows us to obtain more precise saliency clues, laying a solid foundation for the next part of the temporal. In temporal module: we use a set of Encoder-decoder ConvLSTM with different dilated rates, and we use dense connection and skip connection to blend information of different scales. We evaluate our results on four datasets and compare with twelve algorithms. The experimental results show that our algorithm achieved the state-of-the-arts.