Bengali Cyberbullying: Detection, Categorization, and Gender Bias Analysis
Md. Mithun Hossain, Md. Shakil Hossain, Sudipto Chaki, Md. Saifur Rahman, Mohammad Ali Moni · 2024
Cyberbullying has spread like wildfire on social media, seriously harming people’s emotional and physical health. Because of linguistic and cultural differences, addressing this issue in Bengali poses particular difficulties. By emphasizing three crucial tasks—detection, classification, and gender bias analysis—this study offers a thorough strategy to combat cyberbullying in Bengali. To handle multiple tasks concurrently, we provide a unique multi-task learning framework based on a Transformer-based model. The process involves pre-trained tokenization, augmentation, and data cleaning to guarantee that the supplied data is relevant and of high quality. We assess our approach to the investigation of gender bias, the classification of cyberbullying categories, and the identification of particular bullying labels. The Transformer model outperforms baseline models like Bi-LSTM, Bi-GRU, and CNN, with Macro F1 scores of 97.73% for gender classification, 95.31% for categorization, and 98.61% for bullying label detection. These findings demonstrate the model’s efficacy in identifying nuanced linguistic patterns and offer a solid remedy for Bengali cyberbullying analysis, demonstrating its capacity to manage complex text data effectively.