Enhancing Sentiment Analysis with Multimodal Large Language Models
Thresa Jeniffer J, Swetha M, E Raghuvaran, R. N. Ashlin Deepa, R Surendran · 2025
Sentiment analysis is crucial in understanding user opinions, emotions, and attitudes across various domains. Traditional sentiment analysis methods rely primarily on textual data, limiting their ability to capture the full context of human expression, often including multimodal elements such as images, audio, and videos. Existing approaches struggle with ambiguity, sarcasm, and lack of contextual awareness, reducing accuracy and effectiveness. To address these limitations, we propose a novel framework called Sentiment Analysis using Machine Learning with Multimodal Models (SA-ML-MM), which integrates text, images, and audio inputs using large language models (LLMs) enhanced with multimodal capabilities. Our approach leverages ML-based feature extraction and fusion techniques to improve sentiment classification accuracy. The proposed framework is applied in social media analysis, customer feedback interpretation, and emotion recognition. Experimental results demonstrate that SA-ML-MM significantly outperforms traditional text-based models, achieving higher accuracy and robustness in sentiment prediction. By incorporating multiple data modalities, our approach effectively captures nuanced emotional expressions, making sentiment analysis more precise and reliable.