BiasX: “Thinking Slow” in Toxic Content Moderation with Explanations of Implied Social Biases
Yiming Zhang, Sravani Nanduri, Liwei Jiang, Tongshuang Wu, Maarten Sap · 2023
Toxicity annotators and content moderators often default to mental shortcuts when making decisions.This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected.We introduce BIASX, a framework that assists content moderators with free-text explanations of statements' implied social biases, and explore its effectiveness through a large-scale user study.We show that participants indeed benefit substantially from explanations for correctly moderating subtly (non-)toxic content.The quality of explanations is critical: imperfect machine-generated explanations (+2.4% on hard toxic examples) help less compared to expert-written human explanations (+7.2%).Our results showcase the promise of using free-text explanations to encourage more thoughtful toxicity moderation. 1