Impact of Attention on Visual Sentiment Analysis

Parul Goel, Dinesh Kumar Vishwakarma · 2024

With the increasing use of social media among people, the amount of data is increasing, making its analysis more and more challenging. Many existing works focus on textual sentimental analysis. However, with the escalating use of images and videos, the need to develop a robust visual sentiment analysis model has been pressuring more than ever. This work compares the impact of several visual attention models proposed recently: shuffle, triple, and coordinate attention for visual sentimental analysis. Different models focus on different regions and have different generalization techniques used. The models were tested by integrating the attention mechanism with the baseline deep learning model, resnet50. The new models were trained and tested for the popular MVSA-Single dataset. The models were compared on the basis of computational power requirements and the improvement in the performance to find the best possible trade-off. According to experimental results, Coordinate attention has the best accuracy of 89.73%. However, Triple attention has the best trade-off: an accuracy of 0.8932 with 1200 additional trainable parameters in comparison to 530280 in coordinate attention.

Read the paper · More papers on PaperTik