Detection of Hate Speech on Twitter for Arabic Iraqi Dialect Using Stochastic Gradient Classifier

Ahmed Bahaaulddin A. Alwahhab, Vian Sabeeh, Ali Sami Al-Itbi · 2022

People increasingly use social media platforms to communicate and share information worldwide. However, challenges such as verbal misbehaviors and hate messages have been raised and widely disseminated via social media. In the Arabic region, especially in Iraq, Twitter is one of the most prominent social media platforms that has gained popularity. Consequently, an ethical reason behind this research is to build a system that can distinguish Iraqi hate tweets. This work achieved three contributions. First, a dataset of Iraqi tweets has been collected and annotated; this is the first dataset for hate speech in the Iraqi dialect. Second, an algorithm for distinguishing Arabic tweets from other oriental tweets that use the same Arabic glyph was proposed to represent a new addition to the preprocessing steps for the Arabic NLP area. This algorithm was tested and showed efficiency in detecting Urdu about (0.91), Persian (0.73), and Arabic (0.98). Third, the Iraqi hate speech classifier was built using two types of text features depending on the stochastic gradient classifier. The classifier was compared with various machine learning algorithms like SVM logistic regression. The proposed system with stochastic gradient classifier was more efficient than other classifiers achieving precision and recall of 0.80.

Read the paper · More papers on PaperTik