Transformer-based Multimodal Framework for Misogynous Meme Identification

Nitin Kumar Singh, Prativa Das, Pardeep Singh, Satish Chand · IETE Journal of Research · 2025

The recent rise in meme popularity on social media brings both amusement and challenges, with memes being misused to spread misogynistic and toxic content targeting women online. Existing methods for detecting misogyny tend to focus solely on text or visuals, overlooking the need for analyzing multimodal data that combines both images and text. We propose a deep learning framework, namely, the ViDe framework with CROMA-SA module, for automatic misogyny identification in memes. To improve the generalization capability of the proposed framework, we have incorporated state-of-the-art deep learning frameworks, DeBERTa and ViT, to process the text and images in the memes, respectively. The CROMA-SA module, based on cross-modality encoding, attention based late-fusion, and self attention techniques, is proposed to determine the comprehensive context of the meme. To deal with the complex dataset and their class imbalance problem, a loss function named compound loss is introduced, which enhances the model’s ability to learn informative features. To assess the effectiveness and adaptability, we evaluated the proposed ViDe framework on two diversified benchmark datasets: the primary MAMI dataset provided in the SemEval-2022 task 5 and the secondary Hateful Memes dataset. The proposed framework achieved an F1-score of 0.865 and 0.783 on subtasks A and B of the primary dataset, respectively. On the secondary dataset, the proposed framework achieved an AUROC score of 0.814. The experimental findings clearly illustrate that the proposed ViDe framework outperforms existing state-of-the-art multimodal models in both datasets.

Read the paper · More papers on PaperTik