Multilingual multimodal cyberbullying detection through adaptive and hierarchical fusion
Walaa Saber Ismail, Hikmat Ullah, Muhammad Adnan, Farman Ullah · Array · 2026
Detecting cyberbullying in multimodal content (such as memes) is challenging due to complex interactions between images and text, often involving sarcasm, multilingual usage, and other noisy real-world factors. This paper presents a multilingual multimodal cyberbullying detection framework that combines early fusion, late fusion, and hierarchical fusion strategies within a unified architecture. The framework introduces three key modules: Adaptive Cross-Modal Token Integration (ACTI) for iterative early fusion, Context-Adaptive Ensemble with Uncertainty-Aware Gating (CAE-UAG) for dynamic late fusion based on input reliability, and a Hierarchical Contextual Fusion Network (HCFN) that feeds early fused context back into later unimodal processing for refined predictions. Our system leverages state-of-the-art pretrained vision-language models (e.g., CLIP for images and XLM-RoBERTa for text) to learn subtle cross-modal representations (e.g., sarcasm or image–text irony) and uses uncertainty modeling to handle ambiguous or noisy inputs. We evaluate the approach on two benchmark datasets: the English-language Facebook Hateful Memes and the ArMeme dataset of Arabic memes. Experimental results show that our model outperforms multiple baselines (including single-modality models and a strong CLIP-based multimodal baseline), achieving high accuracy, F1-scores, and area under ROC (AUROC) across languages. Notably, it achieves state-of-the-art performance (e.g., 0.85 F1 and 0.88 AUROC on Hateful Memes), surpassing prior fusion methods. The proposed framework represents a significant step toward generalizable, culturally aware, and robust multimodal cyberbullying detection suitable for deployment across diverse social media contexts. • Propose a hierarchical fusion framework that systematically integrates early (ACTI), intermediate hierarchical (HCFN), and late (CAE-UAG) fusion strategies for multimodal cyberbullying detection. • We introduce a novel Context-Adaptive Ensemble with Uncertainty-Aware Gating (CAE-UAG) module that explicitly models prediction uncertainty and incorporates contextual quality signals (OCR confidence, image quality scores, language indicators). • Design a multilingual architecture leveraging XLM-RoBERTa and CLIP that achieves strong cross-lingual performance without language-specific modifications. • We conduct extensive experimental validation including comprehensive ablation studies that quantify each component’s contribution.