FMSA: Few-Shot Multimodal Sentiment Analysis for Social Media via Integrated Prompt Learning and Vision-Language Models
Xianxun Zhu, Heyang Feng, Erik Cambria, Xiaohan Yu, Jose I. Santamaria, Xuhui Fan, Rui Wang · IEEE Transactions on Affective Computing · 2026
Multimodal sentiment analysis (MSA) on social media is increasingly critical for understanding complex emotional expressions, yet it faces significant challenges in data-scarce environments where annotated multimodal content is limited. Here, we present FMSA, a novel framework that integrates prompt-based learning with advanced vision-language models to enable robust few-shot MSA. Our approach leverages an instruction-aware Query Transformer (Q-Former) to dynamically extract and align visual features with task-specific textual prompts, enhancing cross-modal fusion. We introduce a distributed consistency sampling strategy to construct representative few-shot datasets, ensuring statistical diversity under constrained conditions. Evaluated across six benchmark social media datasets-including MVSA-S, MVSA-M, and Twitter-Depression-FMSA outperforms state-of-the-art methods, achieving an accuracy of 63.47% and an F1 score of 57.34% on MVSA-S with full data, and 61.25% accuracy with 56.06% F1 in few-shot settings using just 1% of the data. By fine-tuning lightweight components while preserving pretrained model robustness, FMSA mitigates overfitting and delivers generalizable performance. We also release a curated few-shot dataset as a community resource. This framework advances MSA by offering an efficient, scalable solution for interpreting multimodal emotions in low-resource scenarios.