The Research on Image-text Fusion Emotion Recognition Based on Deep Learning
Xuyang Wang, Xin Zhang · 2024
Today, with the rapid development of short video platforms, people are no longer limited to text through social media expressing emotions, but often express their personal views in the form of multi -modal data, especially the combination of pictures and texts. These multimodal data better reflect the subjective views and emotional tendencies of people. However, in the past, research usually simply combines the data characteristics of different modals. This method does not effectively solve the problem of duplicate and contradictions that may exist between different modes, which affects the accuracy of emotional analysis. In response to the complexity between different modal information and the emotional characteristics of netizens, this study proposes a graphic fusion strategy based on BRA-Transformer. By combining the BERT-BiGRU model, it can more finely extract the characteristics in text data, especially in capturing the context meaning and long-range dependency relationship. In terms of image processing, we use ResNet to extract the image features. Through the attention mechanism and combine the architecture of the Transformer Encoder, it iterates richer edition semantic information. This fusion method can dynamically adjust the weight of graphic modal characteristics, and automatically pay attention to key feature words that are closely classified as emotional classification, effectively integrate graphic information together to improve the accuracy and effect of analysis. Among them, the accuracy and F1 values on the MVSA-Single dataset increased by 1.49% and 1.38%, respectively. The accuracy and F1 values on the MVSA-Multi dataset increased by 1.29% and 1.09%, respectively.