DNMCN: Dual-Stage Normalization Based Modality-Collaborative Fusion Network for Multimodal Sentiment Analysis
Miao Chen, Jin Liu, Xingye Li, Yaohui Zhang, Hongze Liu, Jiajia Jiao, Huihua He · IEEE Transactions on Affective Computing · 2025
Due to the high-quality semantic information provided by the text modality, text-driven models have become the dominant approach for Multimodal Sentiment Analysis (MSA) in recent years. Despite notable progress in previous studies, two primary limitations remain: (i) aligning multimodal features often relies on simple matching of sequence length or feature dimension, which overlooks cross-modal heterogeneity. (ii) existing fusion techniques tend to over-rely on text, potentially diminishing the emotional data contributed by other modalities. To address these issues, in this paper, we propose a Dual-stage Normalization based Modality-Collaborative Fusion Network (DNMCN). Initially, to reduce modality discrepancies, we introduce a dual-stage normalization strategy, where features from different modalities were mapped into a common dimensional space in the first stage to facilitate effective cross-modal comparisons; sequence length inconsistencies caused by cropping and multiscale dimension reduction were addressed in the second stage. Additionally, to achieve high-quality cross-modal mapping without losing non-textual modality information, we propose an Adaptive modality-Collaborative Fusion Transformer (ACF-T) block. Specifically, in ACF-T block, textual semantics are first integrated into any non-text modality via multi-head attention. Next, a novel adaptive weighting strategy is introduced to balance the contribution of fused features and other non-textual modality features, thereby enhancing crossmodal interaction. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches on the public benchmark datasets CH-SIMS, CMU-MOSI and CMU-MOSEI.