News-RO-Offense - A Romanian Offensive Language Dataset and Baseline Models Centered on News Article Comments

Andreea Cojocaru, Andrei Paraschiv, Mihai Dascălu · 2022

The use of offensive language can lead to uncomfortable situations, psychological harm, and in particular cases even to violence.Social networks and websites struggle to reduce the prevalence of these types of messages by using an automated detector.In this paper, we propose a novel Romanian language dataset for offensive message detection.We manually annotated 4,052 comments on a Romanian local news website into one of the following classes: non-offensive, targeted insults, racist, homophobic, and sexist.In addition, we establish a baseline of five automated classifiers, out of which the model based on RoBERT and two layers of CNN achieves the highest performance with an average F1-score of .74.

Read the paper · More papers on PaperTik