Semantic-Aware Visual Saliency Prediction Assisted by Multimodal Large Models
Mohan He, Kaiwei Zhang, Yitong Chen · 2025
Visual saliency prediction aims to simulate the distribution of human attention during observing images, which inherently requires deep semantic understanding of visual scenes. While existing deep learning approaches predominantly depend on models implicitly acquiring high-level semantic features during the training process, which frequently leads to limited perceptual alignment with human perceptual mechanisms. In this paper, we propose a semantic-aware visual saliency prediction pipeline that explicitly integrates deep semantic information assisted by multimodal large models (MLMs). Our pipeline consists of two synergistic components: MLM-based semantic description generation, which extracts textually encoded scene semantics, and multimodal visual saliency prediction network, named CLIPSalNet, which effectively integrates these textual semantics with visual features to produce more accurate and humanaligned saliency maps. Experiments on benchmark datasets demonstrate significant performance improvements, highlighting the effectiveness of introducing multimodal semantic information into saliency prediction tasks.