Misogynistic Content Detection in Roman Urdu Tweets Based on Transformer Models

Waseem Ullah Khan, Salman Ahmed, Safdar Nawaz Khan Marwat, Yousaf Khan, Aftab Khan, Shafi Ullah Kamran · 2024

Facebook, Instagram and Twitter serve as influential social media platforms for individuals where they express and share their thoughts, skills, knowledge and talents with a broad audience. However, these platforms are also used to disseminate offensive content including trolling and content that targets a person’s gender, religion, or race. When women become the targets of such content, it is often manifests as misogyny. In recent years, the growing prevalence of racial and verbal abuse directed at women on social media has drawn considerable attention to the issue of online misogyny and women based offending, by making the automatic detection of such offensive content an urgent priority. Moreover, many researchers have addressed misogyny detection in high resource languages i.e., English, Italian, Arabic, Hindi, and more, tackling this issue in low resource language like Roman Urdu presents a significant challenge. In this paper, a framework is proposed for detection of misogynistic content from Roman Urdu tweets. This paper leverages well known transformer models such as BERT, RoBERTa, GPT-2, and XLNet to train and evaluate misogyny detection performance. The research findings reveal that RoBERTa surpasses the other models in misogyny detection, by achieving the highest accuracy, precision, F1-score and recall i.e., 90.87%, 90%, 89%, and 91%, respectively.

Read the paper · More papers on PaperTik