AMOSI:Attention-Based Multimodal Sentiment Intensity Analysis with LSTM-CNN Architecture for Online Opinion Videos

V V V Bhagya Sree, G Bharathi Mohan, R Prasanna Kumar, M Rithani, Gundala Pallavi · 2024

Through online video sharing websites, people share their thoughts, experiences, and reviews on a daily basis. Researching the subjectivity and sentiment in these opinion videos is becoming more and more popular among academics and industry professionals. Although sentiment analysis works well for text, there isn’t much research done on it when it comes to videos and other multimedia. The largest obstacles to research in this area are the absence of appropriate methodology, baselines, datasets,and statistical analysis of the relationships between data from various modality sources. The Multimodal Opinionlevel Sentiment Intensity dataset (MOSI), the first opinion-level annotated corpus of sentiment and subjectivity analysis in online videos, is presented to the scientific community in this paper. With labels for subjectivity, sentiment intensity, per-frame and per-opinion annotated visual features, the dataset is meticulously annotated. User-generated opinion videos are increasingly prevalent, necessitating advanced sentiment analysis techniques. We present an attention-based LSTM-2D CNN model for multimodal sentiment intensity prediction. Our contributions are three-fold: 1) We collect new annotations of visual gestures and acoustic effects in online videos to augment the MOSI dataset. 2) We design an end-to-end neural architecture to model textual, visual and acoustic modalities. Bi-directional LSTMs capture temporal dependencies in text and audio. 2D CNNs extract spatial visual features. A hierarchical attention mechanism identifies important utterances and frames. 3) Extensive experiments demonstrate 74.7% accuracy, outperforming prior BC-LSTM models by 1.7%. Ablation studies validate the impact of multimodal fusion and attention. Our work provides an improved neural baseline on MOSI, with detailed analysis of multi-view sequential modelling for opinion videos. The additional annotations and strong results will aid future research in this emerging domain.

Read the paper · More papers on PaperTik