Ensemble Learning for Cyberbullying Detection: A Transformer-Based Approach on English and Bangla Texts

Faria Islam, Ahmad Abdullah, Md Yousuf Emon, Fernaz Narin Nur · 2025

Bullying on social media has become a widespread issue in today's digital age as as the anonymity of perpetrators makes it easier to carry out such harmful behavior. Various prevention methods have been introduced to address this problem. Numerous researches have also been conducted to classify cyberbullying content with greater accuracy. The use of transformer models for cyberbullying detection is relatively recent. In this paper, we leverage five transformer models-BERT, RoBERTa, ALBERT, GPT-2, XLNet.-for detecting and classifying English cyberbullying texts, and three models-mBERT, BanglaBERT, and XLM-R-for identifying bullying content in BangIa texts. Finally, we implement a weighted voting ensemble method to enhance accuracy by assigning higher weights to better-performing models. The ensemble approach outperforms the individual models in both English and BangIa datasets, achieving an accuracy of 87.26% and an F1 score of 87.17% in the English, and an accuracy of 87.64% and an F1 score of 87.63% in the BangIa dataset.

Read the paper · More papers on PaperTik