Multi-modal sentiment analysis based on multi-level modal information interaction

Hao Wang, Gulanbaier Tuerhong, Mairidan Wushouer, Xu Guo · Procedia Computer Science · 2025

Aiming to address the issues of insufficient modal interaction, ineffective utilization of the dominant advantage of textual modality, and the loss of modal private features during cross-modal information fusion in existing multimodal sentiment analysis methods, this paper proposes a sentiment analysis method based on multimodal information interaction, termed MLMI. This method enhances sentiment analysis performance by integrating the dominant advantage of textual modality with a multimodal interaction mechanism. The model architecture comprises two core components: 1) the multimodal information interaction module, which employs a dual gating mechanism to facilitate cross-modal information interaction among text, audio, and video modalities. It prevents modal confusion through joint representation space constraints and loss function optimization, thereby promoting effective information fusion while preserving the private information of each modality; and 2) the modal compensation layer, which combines a cosine similarity weighting and scaling mechanism to maintain modal specificity through an independent private encoder. Experimental validation demonstrates that MLMI achieves significant results on the CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets. Specifically, the binary classification accuracy and F1 score on the CMU-MOSI dataset are improved by 1.68% and 1.65%, respectively, compared to the baseline model, highlighting its superior cross-linguistic adaptability and robustness.

Read the paper · More papers on PaperTik