PRISM: Profiling Hate Speech Spreaders using SVM and RoBERTa Models
Suraj Kamal, Ashok Yadav, Vrijendra Singh · 2024
The prevalence of social media in the modern era has led to a concerning issue like the proliferation of hate speech. Twitter, with its massive user base of nearly 237.8 million people actively participating daily, has become a breeding ground for individuals who propagate offensive content, including racism, misogyny, and other forms of derogatory language. Proactive measures to identify and hold accountable individuals spreading hate speech on Twitter are essential. To address this we have proposed work for detection of the hate speech spreader. We have used the dataset provided by PAN @CLEF 2021 and specifically focused on the English language. For model training, we have employed four distinct approaches: Support Vector Machines (SVM),bidirectional LSTM,bidirectional GRU and transformer models. The SVM model employed TF-IDF vectorization to accurately represent the content of the tweets, whereas the transformer model utilized the Roberta tokenizer and a pre-trained Roberta model for effective text encoding. Hyperparameter optimization was performed using random search to enhance model performance. The RoBERTa model outperformed other models with an accuracy of 84%. The SVM model achieved an accuracy of 82%, while the Bidirectional LSTM and Bidirectional GRU models achieved accuracies of 56% and 63% respectively. These results highlight RoBERTa’s superior performance in detecting hate speech spreaders on Twitter.