ConPrompt: Pre-training a Language Model with Machine-Generated Data for Implicit Hate Speech Detection

Youngwook Kim, Shinwoo Park, Youngsoo Namgoong, Yo-Sub Han · 2023

Implicit hate speech detection is a challenging task in text classification since no explicit cues (e.g., swear words) exist in the text.While some pre-trained language models have been developed for hate speech detection, they are not specialized in implicit hate speech.Recently, an implicit hate speech dataset with a massive number of samples has been proposed by controlling machine generation.We propose a pre-training approach, CONPROMPT, to fully leverage such machine-generated data.Specifically, given a machine-generated statement, we use example statements of its origin prompt as positive samples for contrastive learning.Through pre-training with CONPROMPT, we present TOXIGEN-CONPROMPT, a pre-trained language model for implicit hate speech detection.We conduct extensive experiments on several implicit hate speech datasets and show the superior generalization ability of TOXIGEN-CONPROMPT compared to other pre-trained models.Additionally, we empirically show that CONPROMPT is effective in mitigating identity term bias, demonstrating that it not only makes a model more generalizable but also reduces unintended bias.We analyze the representation quality of TOXIGEN-CONPROMPT and show its ability to consider target group and toxicity, which are desirable features in terms of implicit hate speeches.1

Read the paper · More papers on PaperTik