DFNM: Dynamic Fusion Network of Intra- and Inter-modalities for Multimodal Sentiment Analysis
Zebin Li, Junteng Ma, Xia Li, Xin Pan · 2021
Multimodal sentiment analysis aims to leverage multiple modalities involving language, vision and acoustic to analyze the sentiment polarity tendency expressed in text or video. In practice, we usually receive the multimodal information at the same time. For example, when we watch a video, we usually listen to the sound and watch the subtitles in sync. This indicates that different modalities should be synchronously fused in a way of acrossing different timestamps (Inter-modal fusion). Previous studies investigate different methods to obtain the good fusion of multiple modalities by improving the LSTM structure. However, due to the characteristics of LSTM structure, each temporal information of one modality cannot be fully fused with all the temporal information of all other modalities at the same time. Moreover, each temporal information of the same modality also need to be fused (Intra-modal fusion). To fill this gap, in this paper, we design a novel multimodal heterogeneous graph structure to obtain full dynamic fusion of different temporal information from intra- and inter-modalities effectively. Experimental results on public datasets demonstrate the effectiveness of our proposed method.