An Active Learning-Based Alternative Reinforcement Contextual Information Fusion Model for Multimodal Sentiment Analysis

Xianfei He, Yushan Pan, Yangbin Chen, Zuhe Li, Zhijie Xu, Chenguang Yang, Kaiwei Wang · IEEE Transactions on Audio Speech and Language Processing · 2025

Multimodal Sentiment Analysis (MSA) combines data from multiple modalities to accurately interpret human emotional states. Despite recent advancements, deep learning-based MSA models continue to face critical challenges: (1) high costs and resource demands associated with large-scale data annotation, (2) limited effectiveness in capturing global intra-modal context in single-modality feature extraction, and (3) underutilization of complementary cross-modal information in inter-modal interactions. To address these challenges, we propose an innovative model, the Active Learning-based Aternative Reinforcement Contextual Information Fusion (AL-ARCF) model, designed for MSA tasks. Our approach introduces two novel active learning criteria: an abundance criterion, derived from gradient magnitude, and an availability criterion, both aimed at minimizing labeling costs without compromising model performance. Inspired by human learning principles, we also incorporate curriculum learning, gradually increasing learning complexity by dynamically balancing sample difficulty and active learning effectiveness. For enhanced intra-modal context extraction, we propose the Enhanced Global Context Extraction (EGCE) module, which captures detailed spatio-temporal features within unimodal data. To optimize cross-modal interactions, we introduce an Alternative Feature Interaction (AFI) module, leveraging relevance-based feature classification and cross-modal multi-head attention to fully exploit low-level feature correlations across modalities. Extensive experiments on standard MSA benchmarks, the CMU-MOSI,the CMU-MOSEI, and the CH-SIMS datasets, demonstrate that AL-ARCF achieves superior performance compared to existing models, verifying the proposed framework’s effectiveness and robustness.

Read the paper · More papers on PaperTik