Multimodal image-text sentiment analysis based on image attention enhancement: Leveraging CBAM to enhance image features for capturing sentiment information in images for facilitating multimodal sentiment analysis

Xiangyu Kong, Xingxing Yao, Hai Liu · 2023

Traditional sentiment analysis methods are often confined to single modality, failing to fully exploit the information in multimodal data. They overlook the interaction and richness of information among different modalities, relying solely on a single modality to extract emotional features from the data can easily lead to ambiguous results. To address this issue, this paper proposes a multimodal image-text sentiment analysis model based on image attention enhancement, MITSIAE. The model employs Bidirectional Encoder Representations from Transformers (BERT) model for text feature extraction, followed by Visual Geometry Group 19 (VGG19) model for image feature extraction. To enhance the representational capacity of image features, the Convolutional Block Attention Module (CBAM) was employed, which attentively weights image features with channel attention and spatial attention. This enables the model to more effectively capture key regions and features within the images. The enhanced image features are fused with text features, and conduct sentiment classification on the fused image-text features. The proposed approach's effectiveness in the task of multimodal sentiment analysis, which involves both text and images, is demonstrated by the experimental results obtained on the MVSA dataset.

Read the paper · More papers on PaperTik