Data-Efficient Methods For Improving Hate Speech Detection
Sumegh Roychowdhury, Vikram Gupta · 2023
Scarcity of large-scale datasets, especially for resource-impoverished languages encouraged exploration of data-efficient methods for hate speech detection.In this work, we progress implicit and explicit hate speech detection using an input-level data augmentation technique, task reformulation using entailment and crosslearning across five languages.Our proposed data augmentation technique EasyMixup, improves the F1 performance across languages by 0.5-9%.We also observe substantial F1 gains of 1-8% by reformulating hate speech detection as Entailment-style problem.We further probe the contextual models and observe that higher layers encode implicit hate while lower layers focus on explicit hate, highlighting the importance of token-level understanding for explicit and context-level for implicit hate speech detection.1