Objects Affect Where We Look: Improving Human Attention Prediction Using Extra Data
Shuo Zhang, Huike Guo · 2023
Saliency Prediction aims to predict the attention distribution of human eyes given an RGB image, which has been widely considered in human-computer interaction. Most of the recent state-of-the-art methods are based on deep image feature representations from traditional CNNs. However, the traditional convolution could not capture the global features of the image well due to its small kernel size. Besides, the high-level factors which closely correlate to human visual perception, e.g., objects, color, light, etc., are not considered. Inspired by these, we propose a Transformer-based method with object segmentation as another learning objective. More global cues of the image could be captured by Transformer. In addition, simultaneously learning the object segmentation simulates the human visual causing factor. We build an extra decoder for the subtask and the multiple tasks share the same Transformer encoder, forcing it to learn from multiple feature spaces. We find in practice simply adding the subtask might confuse the main task learning, hence Multi-task Attention Module is proposed to deal with the feature interaction between the multiple learning targets. Our method boosts the performances on SALICON and CAT2000 benchmarks compared to other methods, which are two famous datasets for human attention prediction.