Can GPT-4 detect subcategories of hatred?
Raza Ul Mustafa, Noman Ashraf, Nathalie Japkowicz · 2024
Automatic hate speech detection is an important topic of investigation, though simply identifying a social media post as hate-laden may not be sufficient. To properly catalog the kind of hatred and, eventually, generate appropriate interventions, it is important to understand what the hatred relates to. Tropes are defined as significant recurring themes in discourse. Hate speech often revolves around tropes specific to the kind of hatred expressed in the post. Common tropes within Islamophobia and antisemitism, for example, include “Oppression to women” and “Domination and Control”, respectively. The purpose of this paper is to investigate how well the powerful GPT-4 Large Language Model (LLM) can identify the trope expressed in a hateful social media post. In particular, we investigate the performance of GPT-4 on the problem of trope classification in Islamophobia and antisemitism. The results suggest that GPT-4 still presents significant gaps in understanding human text in the context of hate speech, opening the door to future research in that area. Warning: some of the Islamophobic and antisemitic examples and terms shown in this article are of a disturbing nature, but must be included to fully illustrate the purpose of our study.