Annotating hate speech: Three schemes at comparison
Fabio Poletto, Valerio Basile, Cristina Bosco, Viviana Patti, Marco Antonio Stranisci · Institutional Research Information System University of Turin (University of Turin) · 2019
Annotated data are essential to train and benchmark NLP systems.The reliability of the annotation, i.e. low interannotator disagreement, is a key factor, especially when dealing with highly subjective phenomena occurring in human language.Hate speech (HS), in particular, is intrinsically nuanced and hard to fit in any fixed scale, therefore crisp classification schemes for its annotation often show their limits.We test three annotation schemes on a corpus of HS, in order to produce more reliable data.While rating scales and best-worst-scaling are more expensive strategies for annotation, our experimental results suggest that they are worth implementing in a HS detection perspective.1