Fusing Multimodal Sentiment Classification with Improved Deep Multi-View Attentive Network in Image and Text Data Analysis
Pilita A. Amahan · 2024
The use of multimodal sentiment analysis is now gaining its popularity to analyze a user's emotions and feelings in a social media platform. However, the correlations between visual and textual content have been neglected leading to disparity of results. Recently, many of the articles in relation to the field of study concentrate on the analysis of unimodal concept and when there is the study of multimodal analysis it raises the issues of error due to the concern of heterogeneous description between text and image data. Motivated by this status quo, this paper aims to propose a novel improved deep multi-view attentive network in image and text data. The study applies three core processed for: data feature extraction; training, validating, and testing data; and the interpretation of the classified multimodal sentiments. The novelty of the approach is shown in the first and last phase of the study. The initial phase of the study does not only produce a classified sentiment but it also produced sub-classifications to remove heterogeneous information for image and text data. Secondly, image data includes subsets of image sizes to optimize results from different layers and regions. In this, the study achieved 93% of accuracy and though it's a bit lower from the other studies related into, it shows less concern with overfitting of results that is based on the heterogeneous description between image and text data. The process and its analysis demonstrate a superior performance and could be used as a current state-of-the-art technique to evaluate fused multimodal sentiments in social media.