A Hybrid LSTM-CNN Approach for Multimodal Sentiment Analysis: Combining Text and Image Features

Zannirah Muhammed Sammani, Mohammed Abo Rizka · International Journal of Computer Applications · 2025

An efficient deep learning framework is proposed for sentiment analysis that leverages both textual and visual modalities.The architecture integrates Long Short-Term Memory (LSTM) networks for capturing sequential dependencies in textual data with Convolutional Neural Networks (CNNs) for analyzing visual content.This multimodal fusion enhances sentiment classification accuracy.The model is assessed on two benchmark datasets-Memes and MVSA-and its performance is compared to traditional machine learning models such as Support Vector Machines and Logistic Regression, as well as the transformer-based VisualBERT.Although VisualBERT achieves slightly higher accuracy (83.18% on Memes and 81.29% on MVSA), the proposed approach delivers comparable results (77.70% and 80.42%, respectively) while maintaining a much lower computational footprint.This balance between performance and efficiency highlights the model's practical value for applications where computational resources are limited or real-time analysis is required.

Read the paper · More papers on PaperTik