Targeted Identity Group Prediction in Hate Speech Corpora
Pratik Sachdeva, Renata Barreto, Claudia von Vacano, Chris J. Kennedy · 2022
The past decade has seen an abundance of work seeking to detect, characterize, and measure online hate speech.A related, but less studied problem, is the specification of identity groups targeted by that hate speech.Predictive accuracy on this task can supplement additional analyses beyond hate speech detection, motivating its study.Using the Measuring Hate Speech corpus, which provided annotations for targeted identity groups on roughly 50,000 social media comments, we create neural network models to perform multi-label binary prediction of identity groups targeted by a social media comment.Specifically, we study 8 broad identity groups and 12 identity sub-groups within race and gender identity.We find that these networks exhibited good predictive performance, achieving ROC AUCs of greater than 0.9 and PR AUCs of greater than 0.7 on several identity groups.At the same time, we find performance suffered on identity groups less represented in the dataset.We validate model performance on the HateCheck and Gab Hate Corpora, finding that predictive performance generalizes in most settings.We additionally examine the performance of the model on comments targeting multiple identity groups.Lastly, we discuss issues with a standardized conceptualization of a "target" in hate speech corpora, and its relation to intersectionality.Our results demonstrate the feasibility of simultaneously detecting a broad range of targeted groups in social media comments, and offer suggestions for future work on modeling and dataset annotation for this task.