Generative AI for Hate Speech Detection: Evaluation and Findings
Sagi Pendzel, Tomer Wullach, Amir Adler, Einat Minkov · Auerbach Publications eBooks · 2024
Hate speech refers to the expression of hateful or violent attitudes based on group affiliation such as race, nationality, religion, or sexual orientation. In light of the increasing prevalence of hate speech on social media, there is a pressing need to develop automatic methods that detect hate speech manifestation at scale ( Fortuna & Nunes, 2018 ). Automatic methods of natural language processing in general, and hate speech detection in particular, rely heavily on relevant datasets. While researchers have collected several datasets that contain hate speech samples, those resources are scarce. Furthermore, the difficulty in identifying hate speech on social media has led to the use of biased data sampling techniques, focusing on a specific subset of hateful terms or accounts. Consequently, relevant available datasets are limited in size, highly imbalanced, and exhibit topical and lexical biases. Several recent works have indicated these shortcomings and shown that classification m odels trained on those datasets merely memorize keywords, where this results in poor generalization (Wiegand, et al., 2019; Kennedy et al., 2020 ).