A Data-oriented Approach for Detecting offensive Language in Arabic Tweets

Eshrag Ali Refaee · 2021

The growing popularity of social media (SM) platforms has made these platforms a crucial part of modern societies. Users from different cultures, backgrounds, demographics get aboard in an increasing manner to express their views, stances, and opinions on a varied range of topics. Since users on SM can easily hide their real identity, a closer look at daily posts on social medial platforms shows that users do not seem to reflect only their stances and views, but also, they get an opportunity for revealing their behaviors, which could be negative towards the others. Although only a small population of SM users can show negative behavior towards other individuals, groups, and society in general, the impact could be catastrophic. This has resulted in the emerge of terms like cyberbullying, online extremism/hatred/threatening, online trolling, online political-polarity discourse. To ensure safe social networking, the domain of automatic detection of offensive/hatred language has lately grown notably. This work focuses on utilizing a publicly available dataset of Arabic tweets labeled for offensive/non-offensive language. Unlike previous work which focuses merely on developing and tuning machine learning models to be as accurate as possible on the benchmark dataset used, we turn to focus on the characteristics of the offensive language used in SM. The purpose is to have an in-depth look into the dataset to disclose what seems to be hidden patterns in offensive language expressed daily online. Our findings reveal the benefit of using larger training dataset that covers a wide range of offensive language patterns to build robust machine learning classifiers with a better ability to generalize well on highly sparse data used in SM.

Read the paper · More papers on PaperTik