Hate Speech Detection and Bias in Supervised Text Classification
Thomas Davidson · Oxford University Press eBooks · 2023
Abstract This chapter introduces hate speech detection, describing the process of developing annotated datasets, training machine learning models, and outlining the cutting-edge techniques used to improve these methods. The case of hate speech detection provides valuable insight into the promises and pitfalls of supervised text classification for sociological analyses. A key theme running through the chapter is that ostensibly minor methodological decisions—particularly those related to the development of training datasets—can have profound downstream impacts. In particular, the chapter examines racial bias in these systems, explaining why models intended to detect hate speech can discriminate against the groups they are designed to protect and discussing efforts to mitigate these problems. It argues that hate speech detection and other forms of content moderation should be an important topic of sociological inquiry as platforms increasingly use these tools to govern speech on a global scale.