Reducing the Emotional Distress of Content Moderators through LLM-based Target Substitution in Implicit and Explicit Hate-Speech

Nazanin Jafari, James Allan · 2025

Hate speech is often subtle and context-dependent, making it especially difficult to detect particularly when it requires contextual familiarity related to the targeted group.Exposure to hate speech and toxic content can lead to significant psychological harm, including increased stress and anxiety levels and content moderators are particularly vulnerable due to exposure to such harmful material.This work explores the role of personalization in content moderation by examining how alignment between a moderator's background and the targeted group affects emotional and cognitive responses.We propose a target substitution method that replaces references to real communities in hate speech with fictional characters, aiming to reduce emotional distress while preserving the semantic integrity necessary for accurate moderation.Through both automated and human evaluations, we find that substitution significantly reduces emotional distress across all groups with a trade-off in accuracy.Moreover, we observe that moderators demonstrate higher accuracy when moderating content aligned with their own demographic background, even after substitution.This suggests the key role of contextual familiarity in interpreting implicit hate.Additionally, our study highlights the cumulative impact of prolonged exposure to hate speech, showing that moderators experience increased emotional distress over time, particularly in non-targeted scenarios.Despite this, target substitution consistently mitigates distress while maintaining moderation efficacy.

Read the paper · More papers on PaperTik