Cyberbullying Detection in Bengali Social Media Using TF-IDF and Supervised Machine Learning Techniques

Shihab Hossain Shuvo, Abdul Al Mahmud Riaz, Masud Rana Paban, Sajib Howlader, Ratul Bhattacharjee, Md Abu Talha · 2025

Cyberbullying detection in regional languages remains a critical challenge due to limited resources and linguistic complexities. This study focuses on developing an automated framework for identifying cyberbullying content in Bengali social media texts. The research involves collecting a large dataset of Bengali posts, followed by extensive preprocessing steps including text cleaning, tokenization, stopword removal, and stemming to prepare the data for analysis. Feature extraction is performed using Term Frequency-Inverse Document Frequency (TF-IDF) vectorization to convert textual data into meaningful numerical representations. Multiple supervised machine learning classifiers—Decision Tree, Logistic Regression, Random Forest, and Support Vector Machine—are then trained and evaluated to determine their effectiveness in binary classification of bullying versus nonbullying content. The comparative analysis highlights the strengths and limitations of each model, offering insights into the practical applicability of different algorithms for this task. The findings establish a foundation for future advancements in natural language processing tools aimed at fostering safer online environments for Bengali-speaking communities.

Read the paper · More papers on PaperTik