UniTextFusion: A low-resource framework for Arabic multimodal sentiment analysis using early fusion and LoRA-tuned language models
Salma Khaled, Walaa Medhat, Ensaf Hussein Mohamed · Ain Shams Engineering Journal · 2025
Multimodal Sentiment Analysis (MuSA) seeks to interpret human emotions by combining textual, auditory, and visual cues. While this field has advanced significantly in English, Arabic MuSA remains underdeveloped due to limited large language models (LLMs), scarce annotated datasets, dialectal variation, and the complexity of fusing multiple modalities. Cultural elements such as sarcasm and emotional nuance are particularly difficult to capture without multimodal context. An early fusion approach, UniTextFusion, is introduced as a means of overcoming these challenges. This fusion strategy transforms audio and visual inputs into descriptive text, allowing seamless integration with Arabic-compatible LLMs. We apply parameter-efficient Low-Rank Adaption (LoRA) fine-tuning to two generative models—LLaMA 3.1-8B Instruct and SILMA AI 9B. Experiments on our Arabic MuSA dataset show that UniTextFusion improves sentiment classification performance by up to 34% in F1-score over strong unimodal and multimodal baselines, reaching 68% with LLaMA and 71% with SILMA. These results validate our hypothesis that modality textualization combined with lightweight fine-tuning is effective for Arabic MuSA and offers a scalable solution for sentiment analysis in low-resource settings.