TOCAB: A Dataset for Chinese Abusive Language Processing
I Chung, Chuan‐Jie Lin · 2021
This paper introduced TOCAB, a larger dataset for Chinese abusive language detection and classification. This dataset contains 121,344 real sentences collected from a social media site. Several baseline systems built by machine learning or deep learning were proposed to test this benchmark. BERT is the best baseline system which achieves F1-scores of 0.886 in detection and 0.781 in classification. The bootstrap aggregating BERT model, a state-of-the-art system, outperforms our BERT baseline system, with F1-scores of 0.893 in detection and 0.782 in classification.