Towards Safer Communities: Detecting Aggression and Offensive Language in Code-Mixed Tweets to Combat Cyberbullying

Nazia Nafis, Diptesh Kanojia, Naveen Kumar Saini, Rudra Murthy · 2023

Cyberbullying is a serious societal issue widespread on various channels and platforms, particularly social networking sites.Such platforms have proven to be exceptionally fertile grounds for such behavior.The dearth of highquality training data for multilingual and lowresource scenarios, data that can accurately capture the nuances of social media conversations, often poses a roadblock to this task.This paper attempts to tackle cyberbullying, specifically its two most common manifestationsaggression and offensiveness.We present a novel, manually annotated dataset of a total of 10, 000 English and Hindi-English codemixed tweets, manually annotated for aggression detection and offensive language detection tasks 1 .Our annotations are supported by interannotator agreement scores of 0.67 and 0.74 for the two tasks, indicating substantial agreement.We perform comprehensive fine-tuning of pre-trained language models (PTLMs) using this dataset to check its efficacy.Our challenging test sets show that the best models achieve macro F1-scores of 67.87 and 65.45 on the two tasks, respectively.Further, we perform cross-dataset transfer learning to benchmark our dataset against existing aggression and offensive language datasets.We also present a detailed quantitative and qualitative analysis of errors in prediction, and with this paper, we publicly release the novel dataset, code, and models.On the other hand, offensiveness has been described as any word or string of words which has or can have a negative impact on the sense of self or well-being of those who encounter it (Molek-Kozakowska, 2022) -that is, it makes or can make them feel mildly or extremely discomfited, insulted, hurt or frightened.Motivation The dearth of manually-annotated datasets for the tasks of aggression detection and offensive language detection, especially in the Hindi-English code-mixed setting, necessitated us to work in this area.This paper investigates the tasks of aggression detection and offensive language detection on Twitter data.We curate politically-themed tweets and perform manual annotation to create a dataset for the tasks.Our annotation schema is in line with the existing aggressive and offensive language detection datasets.With the help of pre-trained language models, we fine-tune pre-trained language models for both tasks and discuss the obtained results regarding precision, recall, and macro F1-scores.The key contributions of this work are:

Read the paper · More papers on PaperTik