Offensive Language Detection in Social Media for Turkish Language

Ayçe Nida Acar, Sevinç İlhan Omurca · 2023

It is known that the use of offensive language in social media posts has increased considerably in recent years. In this study, it is aimed to determine the use of offensive language in Turkish language. A comprehensive and original data set was obtained by combining two data sets previously used in the literature. There are 31277 data in the first data set (Offenseval-2020) and 35285 data in the second data set (A Corpus of Turkish Offensive Language-troff). The first data set, the second data set and the original data set created from them were preprocessed separately. Afterwards, three data sets were trained in the model created by the LSTM method. The model was tested with 20% of the data set. The experimental results of the study were measured using metrics such as accuracy, precision score, F1 score, and sensitivity score. With the new data set created, higher success has been achieved compared to other data sets. With these successful results, a contribution has been made to the literature in order to determine the use of offensive language in the Turkish language.

Read the paper · More papers on PaperTik