Towards Automatic Generation of Messages Countering Online Hate Speech and Microaggressions

Mana Ashida, Mamoru Komachi · 2022

Warning: This paper discusses and contains content that may be deemed offensive or upsetting.With the widespread use of social media, online hate is increasing, and microaggressions, unintentional offensive remarks in everyday life (Sue et al., 2007), are receiving attention.We explore the possibility of using pre-trained language models to automatically generate messages that combat the associated offensive texts.Specifically, we focus on using prompting to steer model generation as it requires less data and computation than fine-tuning and shows the potential for using prompting in the proposed generation task.We also propose a human evaluation perspective; offensiveness, stance, and informativeness.After obtaining 306 counterspeech and 42 micro intervention messages generated by GPT-2, textscGPT-Neo, and textscGPT-3, we conducted a human evaluation using Amazon Mechanical Turk and found that GPT-3 produces messages of the highest quality among three systems.Also, We discuss the pros and cons of using our evaluation perspectives.We release a corpus of countering hate speech and microaggressions (CHASM), annotated machine-generated counternarratives along with the annotation to promote further research on automatic counternarrative generation and its evaluation.

Read the paper · More papers on PaperTik