Video Saliency Prediction via Joint Discrimination and Local Consistency

Zheng Wang, Ziqi Zhou, Huchuan Lu, Qinghua Hu, Jianmin Jiang · IEEE Transactions on Cybernetics · 2020

While saliency detection on static images has been widely studied, the research on video saliency detection is still in an early stage and requires more efforts due to the challenge to bring both local and global consistency of salient objects into full consideration. In this article, we propose a novel dynamic saliency network based on both local consistency and global discriminations, via which semantic features across video frames are simultaneously extracted and a recurrent feature optimization structure is designed to further enhance its performances. To ensure that the generated dynamic salient map is more concentrated, we design a lightweight discriminator with a local consistency loss LC to identify subtle differences between predicted maps and ground truths. As a result, the proposed network can be further stimulated to produce more realistic saliency maps with smoother boundaries and simpler layer transitions. The added LC loss forces the network to pay more attention to the local consistency between continuous saliency maps. Both qualitative and quantitative experiments are carried out on three large datasets, and the results demonstrate that our proposed network not only achieves improved performances but also shows good robustness.

Read the paper · More papers on PaperTik