Image Description and Aspect-Aware Denoising for Aspect-Based Multimodal Sentiment Analysis

Jiachang Sun, Xiuhong Li · 2025

Aspect-Based Multimodal Sentiment Analysis (ABMSA) aims to combine vision and language to analyze the sentiment of aspect entities in digital content. Many existing methods exploit aspects to fuse visual and textual modalities for multimodal sentiment prediction. However, these methods still have some disadvantages: (1) There are different representation gaps between the visual and image modalities, and these gaps increase the difficulty of alignment between the modalities. (2) When the visual modality is uncorrelated with the textual modality, the visual modality may not enrich the textual modality, which can lead to the introduction of noise and make prediction more difficult. To address these problems, we propose a novel ABMSA architecture model to convert visual modalities and filter noise. Specifically, we proposed an image description (IGD) module to convert the visual modality into the textual modality, thereby avoiding the impact of alignment between different modalities. In addition, we also proposed an aspect-aware denoising (AAD) module, which uses aspects for guidance and enhances the textual modality to denoise the content obtained by the image description module, thereby reducing the noise brought by the visual modality. Experiments on three datasets show that the proposed model has good performance. In addition, comprehensive experimental analysis shows that the proposed model is robust and efficient.

Read the paper · More papers on PaperTik