A Multi-modal Fusion-based Sentiment Analysis Model for Short Videos

Asher Prescott, Juniper Throne, Sullivan Callahan, Matilda Harper · Research Square · 2023

Abstract It is in this information era, leading to the majority of users increasingly like to publish their views online and express their feelings through the form of text pictures and videos, such as Facebook, Instagram, and other social platforms used by a large number of users every day to record their daily life, and through these recorded documents can well reflect the mood of the time. By analyzing these data, researchers can gain insight into users' favorite preferences, and according to their preferences, they can provide precise services to users, thus improving their satisfaction. The traditional approach is limited to linear combinations of sentiment words and rough sentiment classification, to solve the problem of machine learning classification methods. In the deep learning-based image emotion classification method, let the network achieves the method of extracting the significant features of the image by combining the underlying features to form a higher-level feature representation with a more abstract nature. For short videos with complex semantic information and single modal video sentiment analysis often cannot fully express the sentiment tendency of the video. In this paper, we propose a short video sentiment analysis method that incorporates every modal. In the text modality, the text modality data are firstly preprocessed by using the word separation tool, then the word vector corresponding to the text is obtained by using the word embedding tool, and the obtained word vector is used to extract the sentiment features of the text information by using the LSTM network with attention mechanism, and finally the sentiment probability of the text information on the three categories of positive, neutral and negative is calculated by the classifier. On the short video modality, the video feature extraction method of 3D residual dense network is used to establish the video sentiment classification model, and the classifier is used to determine the sentiment probability of short video information. The experiments on several datasets validate the effective performance of our model.

Read the paper · More papers on PaperTik