Investigation of Toxicity Detection Models

Mingyi Wang, Fude Cao, Ning Lu · 2025

Toxic language detection is essential for maintaining safe online platforms, yet many existing models remain vulnerable to evasion tactics. Despite strong performance on clean data, these models often overfit to training distributions, fail to generalize across domains, and are highly sensitive to minor textual variations. We evaluate the performance of toxicity detection models using both the original Jigsaw dataset and its perturbed version with common evasion techniques. This setup simulates realistic deployment conditions and exposes critical weaknesses in model robustness and generalization that are often overlooked by evaluations on clean data alone. Most models degrade substantially when exposed to simple perturbations and fail to retain effectiveness on out-of-domain data. Our findings reveal the persistent vulnerabilities in existing toxicity detection systems.

Read the paper · More papers on PaperTik