Transliterated Bengali Comment Classification from Social Media

Abdullah Al Taawab, Lubaba Tasnia, Mondira Dhar, Md Humaion Kabir Mehedi · 2022

In the era of technological advancement, the internet acts as an essential part of our daily life. People express their opinions on social media through different types of comments. In this paper, machine learning (ML) and deep learning (DL) models have been used to classify transliterated Bengali comments. Due to the lack of a large publicly available transliterated Bengali corpus, we have created our own dataset, consisting of 1,300 transliterated Bengali comments, which is publicly available in Mendeley Data. Moreover, we have applied several ML and DL algorithms, e.g., multinomial naive bayes (MNB), logistic regression (LR), linear SVM, decision Tree (DT), AdaBoost, random forest (RF), RBF SVM, gradient boosting, recurrent neural network (RNN), gated recurrent units (GRU), and long short-term memory (LSTM) for classifying comments. We have implemented different feature extraction techniques to compare the results. Among all these algorithms, logistic regression with countVectorizer performed best with 85.76% accuracy and 85.70% F1 score.

Read the paper · More papers on PaperTik