A Machine Learning Approach to Classify Anti-social Bengali Comments on Social Media

Manash Sarker, Md. Forhad Hossain, Fahmida Rahman Liza, Syed Nazmus Sakib, Abdullah Shahid Farooq · 2022

The growth of social media is causing the emergence of hate speech. Email extortion and cyberbullying are on the rise in Bangladesh, along with online sexual harassment of women. In order to prevent these crimes, studies on Bengali comments on social media have become progressively important. However, the requisite datasets are scarce for this kind of study. The motive of this research is to create a dataset of Bangla comments from social platforms and develop a classifier model as well as to detect whether the comments are social or anti-social quickly and efficiently. 2000 comments were gathered from Facebook and YouTube, two prominent platforms for social media. In our study, an artificial neural network model like Gated Recurrent Unit (GRU), and supervised machine learning classifiers like Logistic Regression (LR), Random Forest (RF), Multinomial Naive Bayes (MNB), and Support Vector Machine (SVM) were utilized in our study to distinguish between anti-social and socially acceptable comments. Finally, language models such as unigrams, bigrams, and trigrams have been implemented in our research. To the best of our knowledge, there are no studies regarding the anti-social classification in Bangla language. This work will help to prevent anti-social activities in Bangla community.

Read the paper · More papers on PaperTik