Leveraging Deep Learning for Cyberbullying Detection in Arabic Tweets
Fatima AlShannaq, Sawsan AlHalawani, Heba Al-Jarrah, Mohammad Shehab · 2025
Cyberbullying is a growing concern in the digital age, particularly on social media platforms like X (formerly Twitter). This study presents a deep learning-based approach for detecting cyberbullying in Arabic tweets using the Arabic Cyberbullying Corpus (ArCybC). Leveraging the Arabic-BERT model, we conducted binary and multi-class classification tasks to identify cyberbullying content and its associated categories or domains. Our methodology incorporated advanced data-balancing techniques, including oversampling and undersam-pling, to address class imbalance challenges. The study aims to answer the research questions: (1) How effectively can deep learning models detect cyberbullying in Arabic tweets? (2) How do data balancing techniques impact classification performance? (3) Can Arabic-BERT outperform traditional classifiers in cy-berbullying detection? These questions are addressed through our methodology, where we employ Arabic-BERT, implement rebalancing techniques, and compare our model's performance with baseline classifiers. Experimental results demonstrate significant improvements in F1-scores, with 89.1% achieved in binary classification and 91.6% in category-level classification after oversampling. These findings highlight the effectiveness of fine-tuned deep learning models in tackling the linguistic complexity of Arabic and the nuanced nature of cyberbullying detection. This study provides a foundation for future research in addressing online harassment in Arabic-speaking digital communities.