MemeFusionNet: A Cross-Linguistic Multimodal Model for Identifying Troll Memes

Tipu Sultan, Hadi Aliakbarpour, Walid El‐Shafai, Md Saib, Md. Khairul Bashar Bhuiyan, M. Shariful Islam, Ahmad Taher Azar, Chakib Ben Njima · 2025

The proliferation of troll memes, which exploit textual and visual elements to propagate misinformation and incite negativity, presents a critical challenge for online content moderation. Existing methods often struggle with cross-linguistic generalization, multimodal fusion, and contextual understanding, limiting their effectiveness in multilingual environments. To address these gaps, we propose MemeFusionNet, a transformer-driven multimodal fusion framework that effectively captures the intricate relationships between images and text. MemeFusionNet integrates a cross-modal attention mechanism based on ViLT to enhance contextual awareness and better detect implicit troll content, such as sarcasm and cultural nuances. Our model demonstrates superior performance on Bangla and English meme datasets, achieving $86 \%$ and $96 \%$ accuracy, respectively, outperforming all existing benchmarks. Its scalable architecture ensures robust cross-lingual adaptability, making it well-suited for large-scale, real-time content moderation.

Read the paper · More papers on PaperTik