Integrating Visual and Textual Cues for Sentiment Analysis: Multi-modal Approach

Hanqi Bai · Applied and Computational Engineering · 2025

With the steady growth of social media and online platforms, sentiment analysis has become a critical task to understand public opinion, customer feedback, and social trends. This study investigates multi-modal sentiment analysis by exploring state-of-the-art models such as Bidirectional encoder representations from transformers (BERT), Bootstrapping Language-Image Pretraining (BLIP), Generative Image-to-text Transformer (GIT), and Contrastive Language-Image Pretraining (CLIP) for sentiment classification tasks using both image and text data. The effectiveness of these models on sentiment analysis tasks is evaluated under different configurations such as text-only BERT, image-to-text augmented BERT, and CLIP-based classification. The results show that while BERT achieves 76% accuracy on text-only sentiment analysis, combining text with image-generated descriptions does not significantly improve performance, with accuracy remaining around 74%. On the other hand, CLIP achieves a moderate 62% accuracy using image-text embeddings. While CLIP performs well on multi-modal mapping, it demonstrates challenges in deep semantic understanding compared to BERT’s 0.76 F1 score on the text-only task. These findings highlight the challenges of effectively merging different modalities and point out future directions for improving sentiment analysis in multi-modal settings, enhancing the ability of models to fully understand the semantic content in both images and text.

Read the paper · More papers on PaperTik