MuSeRAN: Multimodal Sentiment Analysis with Recurrent Attention Networks
Omaia Mohammed Al-Omari, Asem Omari, Esam Mohammed Asem Othman, Atheer Ahmed Alrashed · 2024
This paper introduces a novel Multimodal Aspect-Based Sentiment Analysis (MABSA) methodology by creating a Multimodal Sentiment Analysis model utilizing Recurrent Attention Networks (MuSeRAN). The proposed framework uses the synergy between textual and visual data to improve sentiment classification. We use BERT and Bottom-Up Attention (BUA) models as encoders for text and images to guarantee efficient feature extraction. The primary innovation resides in the interactive fusion module, which amalgamates semantic representations at the token level and employs GRU cells to eliminate noise from the image data. Our iterative attention mechanism enhances aspect-specific sentiment features, facilitating progressive learning for improved classification efficacy. We assessed the proposed model using the Twitter-17 dataset with the Twitter-15 dataset, attaining results surpassing leading baselines. Findings demonstrate that multimodal models, including ours, surpass single-modal methods, significantly depending on text-based attributes for precise sentiment prediction. The T-test verifies the enhancements' statistical significance, affirming our model's robustness. This research underscores the importance of multimodal fusion methodologies in enhancing sentiment analysis applications.