Revisiting Hate Speech Benchmarks: From Data Curation to System Deployment

Atharva Kulkarni, Sarah Masud, Vikram Goyal, Tanmoy Chakraborty · 2023

Social media is awash with hateful content, much of which is often veiled with linguistic and topical diversity. The benchmark datasets used for hate speech detection do not account for such divagation as they are predominantly compiled using hate lexicons. However, capturing hate signals becomes challenging in neutrally-seeded malicious content. Thus, designing models and datasets that mimic the real-world variability of hate warrants further investigation.

Read the paper · More papers on PaperTik