MultiLangMemeNet: A Unified Multimodal Approach for Cross-Lingual Meme Sentiment Analysis
Md. Kishor Morol, Shakib Sadat Shanto, Zishan Ahmed, Ahmed Shakib Reza · 2024
This study introduces MultiLangMemeNet, a novel, unified multimodal approach for meme sentiment classification across diverse languages. The proposed model integrates visual and textual components to effectively capture the multimodal nature of memes. Five language datasets-English, Bengali, Chinese, Hindi, and Tamil-were used for the experiments. In every language examined, MultiLangMemeNet performed consistently better than both state-of-the-art multimodal techniques and unimodal baselines. The accuracy gains ranged from 2.46 % to 13.74 %, indicating significant improvements over the top unimodal vision and text models achieved by the model. Furthermore, MultiLangMemeNet surpassed baseline multimodal techniques, achieving accuracy improvements of 6 % in English (61 % vs 55 %), 2.68 % in Bengali (66.02 % vs 63.34 %), 6 % in Chinese (61 % vs 55 %), 4.2 % in Hindi (73.28 % vs 69.08 %), and 2% in Tamil (47% vs 45%) compared to the next best multimodal approach. The study also explored early and late fusion strategies, revealing language-dependent variations in optimal fusion approaches. The findings indicate a significant advancement in multilingual meme sentiment analysis by demon- strating the efficacy of MultiLangMemeNet in capturing the complex interplay between visual and textual components in memes across various linguistic and cultural contexts.