Ferryman at SemEval-2020 Task 12: BERT-Based Model with Advanced Improvement Methods for Multilingual Offensive Language Identification

Weilong Chen, Peng Wang, Jipeng Li, Yuanshuai Zheng, Yan Wang, Yanru Zhang · 2020

Indiscriminately posting offensive remarks on social media may promote the occurrence of negative events such as violence, crime, and hatred.This paper examines different approaches and models for solving offensive tweet classification, which is a part of the OffensEval 2020 competition (Zampieri et al., 2020; Zampieri et al., 2019b).The dataset is Offensive Language Identification Dataset (OLID) (Zampieri et al., 2019a), which draws 14,200 annotated English Tweet comments (Rosenthal et al., 2020).The main challenge of data preprocessing is the unbalanced class distribution, abbreviation, and emoji.To overcome these issues, methods such as hashtag segmentation, abbreviation replacement, and emoji replacement have been adopted for data preprocessing approaches.The main task can be divided into three sub-tasks, and are solved by Term Frequency-Inverse Document Frequency(TF-IDF) vectorizer, Bidirectional Encoder Representation from Transformer (BERT), and Multi-dropout respectively.Meanwhile, we applied different learning rates for different languages and tasks based on BERT and non-BERTmodels in order to obtain better results.Our team Ferryman ranked the 18th, 8th, and 21st with F1-score of 0.91152 on the English Sub-task A, Sub-task B, and Sub-task C, respectively.Furthermore, our team also ranked in the top 20 on the Sub-task A of other languages (C ¸öltekin, 2020;Sigurbergsson and Derczynski, 2020;Mubarak et al., 2020;Pitenis et al., 2020).

Read the paper · More papers on PaperTik