ArCBDs: A Corpus for Cyberbullying Detection of Arabic Tweets

Mona Alrougi, Ghada Alamoudi, hanan salem Algamdi · IJARCCE · 2024

Cyberbullying is a significant problem with the increase in the popularity of social media platforms, especially in Arabic society.Many detection systems for Cyberbullying have been developed for many natural languages to address this issue.However, Arabic is a resource-limited language that lacks high-quality datasets in many computational research areas.This paper aims to bridge this gap by introducing a new dataset of 10,000 Arabic tweets.Each tweet is manually annotated as either bullying or non-bullying.Then, baseline experiments were conducted on the proposed dataset to evaluate the performance of five models: SVM, NB, LR, AraBERTv0.2-Twitter,and CAMeLBERT-Mix.The results showed that the AraBERTv0.2-Twittermodel achieved the best performance, with an accuracy of 90% and an F1-score of 89%.

Read the paper · More papers on PaperTik