Multimodal Sentiment Analysis Using RNN
Swati Kashyap, Nithin Linga, Kartikeya Vinay Deepak Jakkinapalli, Revanth Ganta, Eeshaan Timmanapalli, Yashmit · 2024
The usage of text, audio, video, and mix-mode content on social networking sites to share ideas and viewpoints has significantly increased. These days, sentiment analysis (SA) and emotion detection (ED) of different social networking posts, blogs, and chats are quite helpful and illuminating for getting the right viewpoints on different situations, entities, or characteristics. Many probabilistic and statistical models based on lexical and machine learning techniques have been used to address these problems. The bulk of the literature indicates that the focus was on enhancing modern instruments, techniques, models, and approaches. Recent developments in deep neural networks have led to intensive testing of several deep learning models to enhance accuracy in the tasks stated above. Deep neural networks that are predominantly used for feature extraction from temporal and sequential inputs are known as recurrent neural networks (RNNs), along with their architectural variants, such as Gated Recurrent Units (GRU) and Long Short Term Memory (LSTM). Textual, aural, visual, or any mix of these can be used as input for SA and related tasks. The role of sequential deep neural networks in multimodal data sentiment analysis is scrutinized. In particular, using RNN and its architectural variants, we offer a comprehensive review of the problems, issues, and approaches related to textual, visual, and multimodal SA.