Tackling Visual Redundancy in Multi-Modal Sentiment Analysis with Salient-Driven Image Pre-Processing
Yu Zhang, Zhenjie Zhao, Haojie Lu, Lifei Xiao, Yifan Liu · Advances in transdisciplinary engineering · 2024
With the form of posts shifting from pure text into multi-modality, more and more researchers turn to the sentiment analysis of multi-modal posts. We focus on analyzing the sentiment of posts composed of text and images. In this task, images provide abundant information for accurate analysis but also produce redundant information. Previous works have tried to address the visual redundancy problem caused by modal heterogeneity, but they often overlook the fact that redundant visual elements can pollute the effective information during the encoding process. To alleviate this issue, in this paper, we propose a novel multi-modal sentiment analysis method, including a Salient-driven Image Pre-processing module (SIP) and a Data Augmentation based Self-supervised learning module (DAS). The SIP module will retain the most salient parts of the image during the pre-processing stage, emphasizing the visual information related to events and topics. The introduction of this module mitigates the challenges such as feature redundancy and feature pollution in the image encoding process. The DAS module enhances the robustness of the model by modelling the raw data and the augmented data of a single sample. Our method is tested on three public benchmarks including MVSA-single, MVSA-multiple and HFM. The results demonstrate the competitiveness of our method against previous approaches by achieving state-of-the-art accuracy on MVSA-single and MVSA-multiple.