Turkish Profanity Detection Enhanced by Artificial Intelligence

Anıl Çelik, Burak Yildirim · 2020

In this work, a hybrid artificial intelligence approach to classify a text as clean or profane is proposed. Focus of the proposed approach is to filter profane text shared by people who abuse the anonymous nature of the internet. Independent of any platform, proposed approach only requires a piece of text to work. In the first step, input text gets pre-processed by an iterative cleaning algorithm. Then, processed text is sent to a heuristic or artificial intelligence based approach depending on the word count. Artificial intelligence based approach uses three unique models to process text. Each model independently calculates profanity probability. For the last step, calculations are sent to a judge model to oversee the classification. Finally, either heuristic or artificial intelligence based approach generates a binary response for classification. It is presumed that the proposed approach will be beneficial for Turkish profanity detection on internet platforms.

Read the paper · More papers on PaperTik