A Robust and Linguistically-Aware Hate Speech Detection System for Roman Urdu
Ehtesham Hashmi, Hasnain Ahmad, Muhammad Tayyab Mazhar, Sule Yildirim Yayilgan, Mehtab Afzal, Sarang Shaikh · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025
Social media sites have developed into a common space for individuals to share their concerns and opinions. There is a chance for individuals and organizations to participate in online behavior that breaches accepted social norms because of the preservation of anonymity and the freedom to communicate ideas without restriction. This leads to a rise in the degree and intensity of hate speech in the online environment. Urdu is the national language of Pakistan and is also widely spoken across several other countries, with over 170 million speakers worldwide. This research addresses the detection of hate speech in Roman Urdu, a prevalent language in Asia, where limited resources exist for mitigating hate speech compared to English. Leveraging machine learning, deep learning, ensemble learning, and natural language processing, we developed a system proficient in understanding Roman Urdu language and culture, capable of identifying diverse hate speech manifestations like abusive language, religious hate, sexism, and racism. We expanded the Roman Urdu Hate Speech and Offensive Language Detection dataset to encompass 30,955 instances, incorporating a novel “Racism” category. Our dataset includes various classes of hate speech such as abusive/offensive, religious hate, sexism, and racism, each reflecting distinct patterns of discriminatory language prevalent in Roman Urdu. After executing text pre-processing, we utilized feature extraction techniques such as Bag of Words and Term Frequency-Inverse Document Frequency embeddings. For model building, we employed several supervised machine learning algorithms, including Random Forest, Decision Tree, Multinomial Naive Bayes, Support Vector Machine, and ensemble methods, coupled with K-Fold cross-validation for robust validation. Additionally, unsupervised learning techniques such as the Gaussian Mixture Model and k-means clustering were also implemented. Deep learning approaches, including Bidirectional Encoder Representations from Transformers, Convolutional Neural Networks, Long Short-Term Memory networks, and multilingual BERT, were explored. Among these, mBERT distinguished itself by achieving an impressive accuracy of 92%, notably surpassing the baseline performance.