DStaL : Multilevel Fusion Classifier Based Analysis of Multimodal Sentiments Using Deep Learning Models
Nasheet Tarik, Ashish Jadhav · Computational Intelligence · 2025
ABSTRACT Multimodal data can more vividly and strangely represent consumers' feelings and sentiments than single‐modal content. Users are moving beyond traditional text‐based content on social media and are now more commonly incorporating images and text to convey their experiences and articulate their opinions. Traditional text‐based techniques have given way to multimodal sentiment analysis, which poses a more complex challenge. This work attempts to address the challenge of sentiment analysis in image‐text posts by presenting a refined recognition method that efficiently leverages textual and visual content information. Initially, a single dataset can be created by combining the MVSA‐Single and MVSA‐Multiple datasets. It is gathered to improve data quality by pre‐processing the data with a Gaussian bilateral filter (GBF) for picture features and lemmatization, stemming, and stop word removal for text characteristics. The normalized term‐inverse document frequency (NorTID) model is used to extract text features. The Convolutional VGG‐16 (ConV‐16) model extracts image characteristics. The Dense Stacked Long Short‐Term Memory Network (DStaL) model is used to independently analyze the collected multimodal information. The obtained features are fused together, and the sentiments are effectively classified using the Multilevel Fusion Classifier (MFuse) model. The study illustrates the enhanced efficacy of the suggested strategy by comparing the results to those of standard approaches using several performance measures. For the MVSA‐Single and MVSA‐Multiple datasets, the findings show a 97.3% accuracy, 94.8% precision, 94.8% recall, 97.3% specificity, 92.1% Matthews Correlation Coefficient (MCC), 36.3% Root Mean Squared Error (RMSE), and 0.07% Mean Absolute Error (MAE).