Hierarchical Fusion Network for Multimodal Sentiment Analysis
Yuyao Hua, Ruirui Ji, Sifan Yang, Yi Geng, Wei Gao, Yun Tan · 2024
In the multimodal sentiment analysis task, the role of text is superior to that of non-text modalities. Effectively learning the interaction relationships between modalities to support the text modality while preserving the specific information inherent to each modality to obtain a fusion representation that includes richer sentiment information is a challenging task. To address this issue, a multimodal sentiment analysis model based on the Hierarchical Fusion Network is proposed. The model is text-dominated and employs feature fusion to capture the interaction between modalities. Additionally, it preserves the unique information within each modality through decision fusion. The sentiment information from audio and visual modalities is processed using Gated Linear Unit, which improves the accuracy of sentiment analysis. Experimental results on the CMU-MOSI and CMU-MOSEI datasets demonstrate the effectiveness of the proposed HFN model, outperforming previous approaches in multimodal sentiment analysis.