A Semi-Supervised Approach for Identifying Cyberbullying Text in Bangla Language

Aifa Faruque, Hossen Asiful Mustafa · 2024

Social media has become increasingly popular, and it is now relatively simple to communicate with people online. As a result, there is more hate speech directed at specific individuals on social media. It is vital to create models that can help in automatically recognizing cyberbullying comments, which is becoming more and more prevalent on social media. Large-scale annotated corpora, which are currently uncommon in languages like Bangla, are typically needed for such models. However, it is exceedingly expensive and time-consuming to manually annotate corpora. To address this issue, we employ a semi-supervised self-training approach to replace the time-consuming human annotation process. We utilize Decision Tree, Logistic Regression, Random Forest and Support Vector Machine for classification and perform experiments by combining different n-gram features to compare the effectiveness of these machine learning methods. Experimental results show that the self-training algorithm with SVM as the base estimator yields a 90.57% F1-score.

Read the paper · More papers on PaperTik