Multi-label Bengali Abusive Comments Classification using Problem Transformation Method

Tahsin Tasnia Khan, Abid Hassan, Md. Faysal Ahamed, Samiul Islam · 2023

Sentiment analysis on the Bengali language is limited to binary and multi-class classifications. This paper focuses on the multi-label classification of the Bengali abusive social media comments dataset having 10220 rows of different social media negative comments using the problem transformation method. The texts are categorized into five labels: toxic, threat, obscene, insult, and racism. Three different multi-label approaches: Binary Relevance, Label Powerset, and Classifier Chain combined with three popular machine learning algorithms: Multinomial Naive Bayes, Random Forest, and Logistic Regression are implemented in this research. This study evaluates the performance of these models by applying them to various classification scenarios, considering five-label, four-label, three-label, and two-label classifications independently. Our research outcomes are evaluated with some performance indicators: Accuracy, Precision, Recall, F1-Scores, Confusion matrix, Macro-average, and Hamming score. The combination of Label Powerset and Logistic Regression outperformed the other two classifier-model combinations because of their quick adaptation to the dataset with an accuracy of 88.07% for five label classifications, 90.12% for four label classifications, 90.56% considering three label classifications and 92.67% for two label classification individually. With 88.07% accuracy, the Label Powerset with Logistic Regression approach can be applied to classify five types of sentiment in a single text.

Read the paper · More papers on PaperTik