Visual Saliency Prediction on 360 Degree Images With CNN

Xinlang Chen, Pan Gao, Ran Wei · 2021

With the rapid development of virtual reality techniques, understanding the visual attention in 360° images has attracted tremendous interest. Due to the lack of a sufficient large-scale 360° image saliency database, it is not straightforward to extend the existing saliency prediction methods from traditional 2D images to 360° images. In this paper, an end-to-end model based on Convolutional Neural Network (CNN) is proposed for 360° images saliency prediction. Firstly, we sample the omnidirectional image into equally-sized patches as the input, so as to reduce the error caused by stretching during equirectangular projection. Afterwards, a model containing two networks is proposed, i.e., basenet and refinenet. The basenet is designed to learn the saliency feature from the planar 2D patch, while the refinenet is to refine the saliency from the basenet by incorporating the observer bias towards the equator in spherical space. In addition, a new loss function is proposed, where both the distribution and location based evaluation metric are considered. Finally, we project each pixel of each patch back to the equirectangular image to generate the complete saliency map. Our experiments show that the proposed model outperforms the state-of-the-art method.

Read the paper · More papers on PaperTik