FusionBullyNet: A Robust English—Arabic Cyberbullying Detection Framework Using Heterogeneous Data and Dual-Encoder Transformer Architecture with Attention Fusion
Mohammed A. Mahdi, Muhammad Asad Arshed, Shahzad Ahmed Mumtaz · Mathematics · 2026
Cyberbullying has become a pervasive threat on social media, impacting the safety and wellbeing of users worldwide. Most existing studies focus on monolingual content, limiting their applicability to online environments. This study aims to develop an approach that accurately detects abusive content in bilingual settings. Given the large volume of online content in English and Arabic, we propose a bilingual cyberbullying detection approach designed to deliver efficient, scalable, and robust performance. Several datasets were combined, processed, and augmented before proposing a cyberbullying identification approach. The proposed model (FusionBullyNet) is based on fine-tuning of two transformer models (RoBERTa-base + bert-base-arabertv02-twitter), attention-based fusion, gradually unfreezing the layers, and label smoothing to enhance generalization. The test accuracy of 0.86, F1 scores of 0.83 for bullying and 0.88 for no bullying, and an overall ROC-AUC of 0.929 were achieved with the proposed approach. To assess the robustness of the proposed models, several multilingual models, such as XLM-RoBERTa-Base, Microsoft/mdeberta-v3-base, and google-bert/bert-base-multilingual-cased, were also trained in this study, and all achieved a test accuracy of 0.84. Furthermore, several machine learning models were trained in this study, and Logistic Regression, XGBoost Classifier, and Light GBM Classifier achieved the highest accuracy of 0.82. These results demonstrate that the proposed approach provides a reliable, high-performance solution for cyberbullying detection, contributing to safer online communication environments.