Fuse-MD: A culturally-aware multimodal model for detecting misogyny memes
Rahul Ponnusamy, R. Saranya, Bhuvaneswari Sivagnanam, Anshid Kizhakkeparambil, Dhruv Sharma, Paul Buitelaar, Bharathi Raja Chakravarthi · Natural Language Processing Journal · 2026
Warning: This study involves an analysis of meme content that may contain offensive or upsetting material, used only for research and illustrative purposes. The data presented does not represent the views or convictions of the authors or their associated organizations. The emergence of social media has transformed global communication, enabling cultural expression and the exchange of ideas across platforms such as Facebook, X, and Instagram. This openness has facilitated the dissemination of harmful content, including misogyny, which surprisingly appears in the form of memes, widely circulated online images embedded in text that convey humor or commentary. Misogynistic memes often mock or criticize women’s experiences, perpetuate negative stereotypes, reinforce gender discrimination, and promote violence. Detecting such content is particularly challenging in low-resource languages such as Tamil and Malayalam, due to the limited availability of linguistic resources and tools. This study introduces the Misogyny Detection Meme Dataset (MDMD), the first multimodal dataset specifically curated for detecting misogyny memes in Tamil and Malayalam. We further conducted a shared task using MDMD and established baselines for both languages. This study proposes the Fusion-based Multimodal Framework for Misogyny Meme Detection (Fuse-MD), which employs a transfer learning approach to identify misogynistic memes in low-resource languages. Comparative analysis shows that element-wise fusion performs best for Tamil, while gated fusion is optimal for Malayalam, highlighting the role of cultural and linguistic factors in fusion design. We also introduce a “Threshold Optimization Technique via Macro-F1 Calibration,” which calibrates prediction thresholds on the development set and helps balance the accuracy loss caused by quantization. Experimental results demonstrate that Fuse-MD outperforms the shared task baseline, providing a robust framework for detecting misogyny memes in low-resource languages and laying the foundation for future research.