An Interpretable Machine Learning Approach for Bengali Toxic Comments Detection
Md. Zobayer Ibna Kabir · 2025
Toxic comment detection is critical for providing a positive online environment and to detect this type of comment is very important for people’s mental health. In this research, a machine learning based pipeline was created to detect if a comment is toxic or not. For prediction purposes, two datasets were utilized.The first dataset, which is publicly available, contains a total of 16,073 comments. Out of 16,073, 8,488 comments are classified as toxic, and 7,585 comments are classified as not toxic.The second dataset, also publicly available and collected from Mendeley Data, contains 44,001 comments. Among these, 28,661 comments are classified as toxic or bully, while 15,340 comments are classified as not toxic or not bully. Preprocessing was completed, and features were extracted using TF-IDF, before training the models on the datasets. Then we evaluated the results using a confusion matrix and performance metrics like precision, f1 score, recall, accuracy. For Dataset-1, the Stochastic Gradient Descent (SGD) classifier achieved the best performance among the implemented models, with a weighted F1 score of 0.91 and an accuracy of 91.44%.For Dataset-2, the Stochastic Gradient Descent (SGD) classifier demonstrated the highest performance among the implemented models, achieving a weighted F1 score of 0.88 and an accuracy of 88.42%..To explain the model’s prediction, LIME framework was implemented. With the use of explainable artificial intelligence, the decision-making process of our models becomes more transparent. Source Code: https://github.com/ZobayerAkib/An-Interpretable-Machine-Learning-Approach-for-Bengali-Toxic-Comments-Detection-IEEE2025