Comparative Evaluation of Pre-Trained Models for Hate Speech Detection on Social Media

Mohini Chakarverti, Anurag Goswami, Ashima Yadav · 2024

The growing spread of hate speech across social networking is getting harder to use that demands ever further sophistication in measurement tools. This paper sought to do the same by understanding if currently available pre-trained transformer models can be effective hate speech identifiers while measuring their performance on critical metrics. In particular, the study aimed to do the same for the best available transformer-based transformers - ALBERT, BART, BERT, DistilBERT, and RoBERTa - through a comparison of their F1 Score, F1.5 Score, recall, accuracy, and precision. This made it much easier to understand the model’s strengths and weaknesses. It is RoBERTa that emerges as the top performer with an accuracy of ${0. 8 9 1 2}$ $(89 \%)$. At the same time, it is easy to see that this model can become a game-changing tool for identifying hate speech using existing infrastructure. Additionally, one can also better understand the tradeoffs this model makes between precision and recall. It makes it easier to understand how these models operate when deployed. In the end, this paper shows how it is possible to further the current discussion around AI for good by presenting a relatively unbiased view of current achievement to create safer-online-spaces. As such, this paper also aims to influence how future hate speech detection projects are developed and deployments are managed.

Read the paper · More papers on PaperTik