Detecting Hate speech in Arabic Literahire Tweets

Sara Marhaba, Arwa Abdelrahman, Afnan Y. Alomairi, Hatoon Alazwari, Fatma Heiba, Areej Althubaity · 2022

Middle Eastern people are one of the largest populations on Twitter. In recent years, more Arabic writers, especially young ones, have used Twitter to publish their literary works to a wider audience. This research focused on two literary genres: prose and poetry, to detect hate speech in Arabic literary texts that were published in Twitter platform. Arabic tweets wm be scraped, pre-processed, and classified into hate or non-hate tweets using five different machine learning algorithms; namely: Support Vector Machines, Naive Bayes, Random Forest, Gradient Boosted Decision Trees, and Extra Tree Classmer in three different scenarios: unbalance, undersampling data, and over-sampling data. We compare the performance of these algorithms based on four evaluation metrics; namely: accuracy, precision, recall, and Fl-score. Our results show that the RF algorithm produces 95.45% accuracy, 98.63% precision, 92.17% recall, and 95.29% F1-score; hence, we decided to display RF classified tweets on a website called (ميراث الأدباء), that means Wnters’ Legacy in the English language, to preserve and enable readers to easily access Arabic literature tweets.

Read the paper · More papers on PaperTik