IMCN: Identifying Modal Contribution Network for Multimodal Sentiment Analysis
Qiongan Zhang, Lei Shi, Peiyu Liu, Zhenfang Zhu, Liancheng Xu · 2022 26th International Conference on Pattern Recognition (ICPR) · 2022
Multimodal sentiment analysis (MSA) aims to obtain the emotional polarity of language by analyzing multiple forms of human language, facial expressions, and vocal intonation. The traditional MSA model focuses on the fusion between modalities, ignoring the different contributions of language, visual, and acoustic. Thus different information of modality possesses different importance. To further explore the contributions of different modalities, we propose a highly generalized identifying modal contributions network(IMCN), which contains modality interaction module, modality fusion, and modality joint learning units in the framework. Specifically, we first designed a language modality gain detection module to make reasonable use of visual and acoustic information and reduce the noise of modal information. Secondly, crossmodal attention is used to enrich modal information. Finally, we perform joint learning of unimodal and multimodal modalities to explore the optimal solution for multimodal output. We compared with other popular multimodal sentiment analysis models and obtained better sentiment classification results on CMU-MOSI and CMU-MOSEI benchmark datasets. We also further validated the effectiveness of different modules of IMCN through ablation experiments and discussed the ideas of IMCN design.