Combining Pre-Trained Language Models and Features for Offensive Language Detection

Zhenming Li, Kazutaka Shimada · 2022

Nowadays, people often express their abusive and offensive thoughts to others on social media easier. The abusive and toxic comments hurt others seriously. Therefore those abusive and toxic comments should be detected properly through natural language processing. In this paper, we focus on two types of features in offensive language: word-level and sentence-level fea-tures. We use lexicon-based and standard bag-of-words features as the word level. We introduce BERT-based and DeepMoji-based features as the sentence level. We apply the four features to a machine learning approach: support vector machines. We evaluate the method using the combinations of four features with a dataset, Curious Cat. The best F1 score was generated by the method with all features. This result shows the effectiveness of our proposed method. In addition, the experimental result indicates that DeepMoji generated from Twitter data is better than BERT which is generated from written language, for an offensive language detection task about social media data.

Read the paper · More papers on PaperTik