Evaluating machine learning techniques for detecting offensive and hate speech in Iraqi tweets

Laith F. Jumaa, Aqeel Ali Al-Hilali, Mustafa Bashar, Hussein Alaa Diame, Ali A. Saber · IET conference proceedings. · 2025

Over the last several years, there has been a conspicuous surge in derogatory and inflammatory remarks on Twitter, often entw ined with racial and religious attitudes. This phenomenon is especially noticeable in Iraq, where Arabic is the main language spok en on the site. Social media posts from Iraq provide distinct difficulties because of the convergence of hate speech, foul language, and freedom of expression. Hence, customized techniques and a domain-specific Arabic corpus are crucial for precisely detecting hate speech and inflammatory tweets. While machine learning has shown its effectiveness in other Arabic settings, our work especia lly aims to assess several machine learning approaches tailored to this particular goal. We compiled an Arabic dataset consisting of tweets from Iraq and used sophisticated techniques to extract and evaluate character n-grams, word n-grams, negative emotion, and syntactic-based characteristics. The aforementioned techniques were integrated with hyper-parameter optimization, ensemble, and multi-tier meta-learning models, including support vector machine, logistic regression, random forest, and gradient boosting approaches. This study primarily focuses on well-known machine learning models for text categorization, including CNN, ELMo, and BERT. We used these methodologies on our datasets and, thereafter, merged the classifiers to improve the accuracy of categorization. The findings indicate substantial improvements in both the accuracy of classification and the F1-score.

Read the paper · More papers on PaperTik