Machine Learning Approach to Detection of Offensive Language in Online Communication in Arabic
Azalden Alakrot, Muftah Fraifer, Nikola S. Nikolov · 2021 IEEE 1st International Maghreb Meeting of the Conference on Sciences and Techniques of Automatic Control and Computer Engineering MI-STA · 2021
This paper presents the results of several machine learning experiments, conducted with a dataset of YouTube comments in Arabic. The experiments aim at studying the impact of various text preprocessing, feature-extraction and feature-selection techniques on the accuracy of a document classifier for detection of offensive language in online communication in Arabic. Regarding data pre-processing, our experiments focus on filtering out noisy characters and normalising inconsistencies present in casual online writing in Arabic. The combined effect of these data preprocessing techniques and a few feature-extraction and feature-selection methods is then evaluated by training document classifiers. Our results give evidence that it is possible to train a classifier for the detection of offensive language on Arabic social media with reasonable overall accuracy of 0.84, and precision, recall and F1-score of 0.89, 0.76 and 0.81, respectively.