Performance Evaluation of Pre-Trained CNN Models for Visual Saliency Prediction
Bashir Ghariba, Mohamed Shehata, Peter McGuire · 2020
Human Visual System (HVS) has the ability to focus on specific parts of the scene, rather than the whole scene. This phenomenon is one of the most active research topics in the computer vision and neuroscience fields. Recently, deep learning models have been used for visual saliency prediction. In this paper, we investigate the performance of five state-of-the-art deep neural networks (VGG-16, ResNet-50, Xception, InceptionResNet-v2, and MobileNet-v2) for the task of visual saliency prediction. In this paper, we train five deep learning models over the SALICON dataset and then use the trained models to predict visual saliency maps using four standard datasets, namely: TORONTO, MIT300, MIT1003, and DUT-OMRON. The results indicate that the ResNet-50 model outperforms the other four and provides a visual saliency map that is very close to human performance.