Detection of Abusive Language for YouTube Comments in Urdu and Roman Urdu using CLSTM Model

Vandana Rajput, Sandeep Singh Sikarwar · Procedia Computer Science · 2025

The explosion of abusive language in online user comments poses a growing concern, contributing to the rise of cyberbullying directed at both individuals and specific groups. Automatic detection and analysis of abusive languages from online or internet comments has been extensively explored in the literary work for the English language. Therefore, this study extends these efforts to detection of abused language in the Urdu comments and also in the Roman Urdu comments. By employing a diverse set of machine learning and deep learning models. Convolutional Neural Network (CNN), Bidirectional Long Short-Term Memory (BLSTM) and Character-level LSTM (CLSTM). The research work used these models to a substantial dataset comprising in Urdu comments and a smaller dataset featuring over in Roman Urdu comments. The inherent complexities natural language construct strumming. English-like Roman Urdu characteristics, and the distinct Natalie style of Urdu provides unique obstacles for processing and classifying comments using deep learning techniques as well as machine learning. The experiment results show that the proposed method gives the best result i.e. 96.2% of accuracy for Urdu comments and 91.4% for Roman Urdu comments. These findings contribute to the advancement of cross-script abusive language detection.

Read the paper · More papers on PaperTik