KS@LTH at SemEval-2020 Task 12: Fine-tuning Multi- and Monolingual Transformer Models for Offensive Language Detection
Kasper Socha · 2020
This paper describes the KS@LTH system for SemEval-2020 Task 12 OffensEval2: Multilingual Offensive Language Identification in Social Media.We compare mono-and multilingual models based on fine-tuning pre-trained transformer models for offensive language identification in Arabic, Greek, English and Turkish.For Danish, we explore the possibility of fine-tuning a model pre-trained on a similar language, Swedish, and additionally also cross-lingual training together with English.Overall we find that monolingual models achieve higher macro-averaged F1 score.With cross-lingual training of Danish together with English, we achieve better results than by training on the small Danish dataset alone.For Arabic, Danish, English, Greek, and Turkish, we obtained macro-averaged F1 scores of 0.890, 0.775, 0.916, 0.848, and 0.