A Comparative Study of Cyberbullying Detection in Bangla Using Machine Learning and Deep Learning Algorithms
M. Salman Sheikh, Sumiya Ansari, Md. Shiful Islam Fahad, Kazi A Kalpoma · 2024
Social networks have made communication easier and more convenient, enabling users to share information. But this same platform can also encourage negative interactions, like online abuse and cyberbullying. Identifying and solving cyberbullying is hard. It can cause mental and emotional distress. Young boys and girls are often the primary targets of such harassment through cyberbullying resulting in suicide. To identify bullying language on Facebook, especially in Bangla, this study used Natural Language Processing (NLP) approaches. Here, we implemented several Machine Learning (ML) models, such as Random Forest (RF), Logistic Regression (LR), and K-Nearest Neighbors (KNN), as well as Deep Learning (DL) models such as Long Short-Term Memory (LSTM) and Convolutional Neural Networks (CNN). Additionally, we utilized Transformer-based architectures, including the pre-trained Bidirectional Encoder Representations from Transformers (BERT) model, to identify bullying patterns in online texts. At first, we pre-process the text, then apply seven different feature extraction techniques, and test the accuracy of the models. According to our study, the Random Forest model combined with the FastText embedding technique reached 84% accuracy, and LSTM gained 85% accuracy. Considering all the evaluation parameters, the Random Forest model with the FastText embedding is the best-performing model among the feature extraction-based approaches and DL models. Random Forest with Word2Vec gives almost the same result as Random Forest with FastText.