A gaze prediction model based on Vision Transformer and VGG-16
Yidong Hu, Tong Li, Ying Zeng, Yuanlong Gao, Bin Yan, Zhongrui Li · 2025
Currently, most saliency prediction models are stimulus-driven, and the regions they highlight are the salient areas from the machine vision. However, intelligent agents represented by humans can flexibly adjust their visual focus according to task requirements and only pay attention to task-related targets. Eye gaze information, as a concrete manifestation of human visual focus under specific task-driven conditions, can effectively improve the performance of models in corresponding tasks. In this paper, we integrate gaze information and propose a gaze prediction model based on Vision Transformer and VGG-16. This model is designed to predict the human gaze regions in the cross-view matching task from street-view images to aerial-view images. This model uses Vision Transformer to extract the low-level visual features of images and VGG-16 to extract the high-level semantic features of images respectively. It adopts a multi-level feature fusion strategy to strengthen feature representation. In addition, through gaze metric guidance and residual connections, task-related information is effectively introduced to optimize the model performance. Experimental results show that this model achieves competitive results on both self-constructed datasets and public datasets.