Mitigating Online Harassment: Machine Learning Approaches for Hate Speech Detection in Transliterated Bengali Comments

Tahbib Manzoor, Md. Wahidur Rahman Araf, Monjurul Sharker Omi, Tanvir Ahmed Abir, Arpan Das Abir, Irin Hoque Orchi, Faisal Bin Ashraf · 2023

In the era of widespread online communication, the detection of hate speech has become increasingly critical for maintaining a healthy digital discourse. This significance is magnified when considering languages with unique characteristics, such as transliterated Bengali, where challenges in distinguishing hate speech abound. This work undertakes the task of exploring machine learning algorithms to tackle this challenge and contribute to the broader effort of fostering a respectful and inclusive online environment. The study introduces a novel dataset for hate speech detection in transliterated Bengali text, employing two distinct data preprocessing approaches—TF-IDF and Bag of Words. Eight diverse machine learning algorithms are then applied to evaluate their performance under each preprocessing technique. The results showcase the efficacy of specific algorithms, with Multinomial Naive Bayes excelling in the binary dataset and Logistic Regression emerging as a top performer in the multiclass dataset. Despite encountering challenges like imbalanced data and word length distribution, our models demonstrate enhanced precision and recall. This work not only contributes valuable insights to the field but also provides a new dataset, paving the way for future advancements in hate speech detection and model robustness.

Read the paper · More papers on PaperTik