Automated Bengali abusive text classification: Using Deep Learning Techniques

Showni Rudra Titli, Shimul Paul · 2023

In recent years, the breakdown of most human causes that hampers a large number of people is verbal abuse. The most noticeable fact is that most of the verbal abuse is made online by people behind the screen. While this problem has occurred with growing modern technologies, a modern solution is needed to solve the drawback. The growing rate of abusive comments and hate speech online is rapidly increasing on a large scale, and manual reports and corrections cannot help this critical situation. This proposed model introduces automation of the hate speech filtering process through deep learning. This work uses only Bengali datasets to find the real classification, as many Bengali and English mixed research works are available in this field. Using the intelligent automated classification of text comments in a limited resource-constrained language (example: Bengali) is critical for several reasons. This proposed system classifies 'Personal Offensive Text', 'Geographical Offensive Text', 'Religious Offensive text', 'Crime Offensive Text', 'Entertainment, Sports, Meme Tiktok, and others'. By combining two large datasets for this research and employing Bengali BERT, the best feasible outcome has been achieved.This paper indicates the complete results of the combined two different datasets and classified into nine segments, while the proposed Bengali BERT models achieved the highest accuracy of 0.706 and a weighted f1 score of 0.705 in the identification and classification tasks.

Read the paper · More papers on PaperTik