BRUMS at SemEval-2020 Task 12: Transformer Based Multilingual Offensive Language Identification in Social Media
Tharindu Ranasinghe, Hansi Hettiarachchi · 2020
In this paper, we describe the team BRUMS entry to OffensEval 2: Multilingual Offensive Language Identification in Social Media in SemEval-2020.The OffensEval organizers provided participants with annotated datasets containing posts from social media in Arabic, Danish, English, Greek and Turkish.We present a multilingual deep learning model to identify offensive language in social media.Overall, the approach achieves acceptable evaluation scores, while maintaining flexibility between languages. IntroductionSocial media has become a normal medium of communication for people these days as it provides the convenience of sending messages fast from a variety of devices.Unfortunately, social networks also provide the means for distributing abusive and aggressive content.Given the amount of information generated every day on social media, it is not possible for humans to identify and remove such messages manually, instead it is necessary to employ automatic methods.As offensive language becomes pervasive in social media, scholars and companies have been working on developing systems capable of identifying offensive posts, which can be set aside for human moderation or permanently deleted (Risch and Krestel, 2018).Along with these studies, a few shared tasks have been organized on detecting offense in social media, such as HatEval (Basile et al., 2019), HASOC (Mandl et al., 2019), TRAC (Kumar et al., 2018) and OffensEval (Zampieri et al., 2019b) co-located with SemEval 2019.In semeval 2020, OffensEval returns for the second time with multilingual offensive language identification in social media.OffensEval 2 focuses on several languages Arabic, Danish, English, Greek and Turkish motivating participants to submit multilingual offensive language identification systems.This paper revisits the problem of offensive language identification describing our submission to the SemEval-2020 Task 1 (Zampieri et al., 2020).The remainder of this paper is structured as follows: Section 2 describes the related work done in the field of aggression detection, Section 3 has a description of the dataset.Section 4 describes the system that was submitted, split into a description of how the data was processed and the architectures of the classifiers that were used.Section 5 presents an analysis of the results of our evaluation of the different architectures, as well as of the final submission.Finally, Section