An evaluation of automated methods for hate detection

Oskar Lindholm · KTH Publication Database DiVA (KTH Royal Institute of Technology) · 2019

Derogatory, foul, hateful and/or prejudiced comments or even threats directed at other individuals have become a common phenomenon in many digital environments. This is a problem that effects many levels of society, and being able to battle it is therefore of utmost importance. The large amount of data created every day creates a need for well working automatic methods for detecting this type of content. The subjective na- ture of hate, as well as the diversity of how it can be expressed, however, makes the creation of such methods somewhat difficult. In this thesis three different automated methods, developed by the Swedish defence research agency (FOI), for hate detection in texts have been evaluated. To aid in the evaluation of these methods and the disambiguation of hate as a concept, an attempt at defining hate based on psychology literature has also been made. The methods are tested using two different data sets: one handpicked set of comments aimed to test the variety in each methods hate detecting ability, as well as one in-the-wild-set aimed at testing the methods performances in a scenario of realistic application. The result shows a major difference of performance based on the set the methods are tested on. As well as the possible improvements that can be made to each method and the weaknesses of each approach, the re- sult shows the difficulty of creating reliable methods for automated hate detection in general.

Read the paper · More papers on PaperTik