MESA: A Multimodal Sentiment Analysis Model Integrating Multi-Expert Mechanism and Enhanced Semantic Attention
Xing Lyu, Tianle Yu, Lihua Huang, Yinping Zhang, Jingwei Zhang · 2025
With the increasing prevalence of user-generated content on social media platforms, multimodal sentiment analysis (MSA) has emerged as a critical research area, aiming to jointly interpret textual and visual information to understand users' emotional states. However, existing fusion strategies often struggle with modality misalignment and redundant feature representa-tions, which degrade model performance. To overcome these challenges, this paper proposes MESA, a multimodal sentiment analysis model that integrates an Enhanced Semantic Attention (ESA) module and a Multi-Expert fusion mechanism guided by uncertainty. Built upon BERT for text and ResNet50 for images, MESA enhances semantic alignment across modalities and dynamically fuses modality-specific features based on uncertainty cues. Experiments on MVSA-Single and MVSA-Multiple datasets show that MESA significantly outperforms several strong baselines in terms of Accuracy and F1-Avg, demonstrating its effectiveness and robustness.