Multimodal emotion recognition based on long-distance modeling and multi-source data fusion
Jia Cai, Yuhui Lin, Ruihong Ma · 2023
With the prevalence of social media, people are increasingly inclined to express emotions in various forms such as text, images, audio, and video, which has made emotion recognition tasks more complex and challenging. Although traditional deep learning algorithms such as CNN and RNN have been widely used in multimodal emotion recognition, there are still challenges in fusing multimodal features and handling long-distance dependencies which affect the accuracy of emotion classification. To address these issues, this paper proposes a multimodal emotion recognition model based on long-distance modeling and multi-source data fusion. The model uses Transformer to encode concatenated multimodal feature vectors, and uses multi-head attention mechanism for feature fusion. It also introduces cross-layer residual connections to enhance the model's generalization ability. Finally, a multilayer perceptron is used to learn the complex relationship between feature vectors and achieve more accurate emotion classification results. The results on the IEMOCAP dataset show that the model's average accuracy reaches 59.3% and F1 value reaches 56.8%. Compared with the baseline methods, the proposed model improves the accuracy, demonstrating its effectiveness in emotion recognition tasks.