Cost-Sensitive Learning and Ensemble BERT for Identifying and Categorizing Offensive Language in Social Media

Fajar Muslim, Ayu Purwarianti, Fariska Zakhralativa Ruskanda · 2021

Some people often abuse freedom of expression on social media to carry out offensive actions. So we need mechanisms to keep social media conducive. This research aims to identify and categorize offensive language on social media, which consists of three subtasks: Offensive language identification (subtask A), Automatic categorization of offense types (subtask B), and Offense target identification (subtask C). These three subtasks use the OLID dataset (Offensive Language Identification Dataset) [1]. The previous research which utilized fine-tuning BERT achieved competitive performance at the OLID dataset (Zampieri. et al. 2019). In this research, we improve fine-tuning BERT performance using cost-sensitive learning and ensemble technique. The evaluation results on test data beat state of the art on subtask B (F1-score 0.7776), second position on subtask A (F1-score 0.8207), and second position on subtask C (F1-score 0.6574).

Read the paper · More papers on PaperTik