A Multi-Output BERT Framework for Abusive Comment Detection and Sentiment Analysis on Low-Resource Language

Mansi Yagnik, M. Hashmi, Deepika Kumar, Khushi Jain, Ekagrah Grover, D. Jude Hemanth · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025

In the modern digital world, social media has become essential for interpersonal interaction by promoting the interchange of ideas and points of view. But there are difficulties in this digital environment, especially concerning rude behavior and offensive remarks. To address both problems at once, the research focuses on sentiment analysis and abusive comment detection in social media interactions. The dataset contains Hate Speech and Offensive Content Identification (HASOC) data from 2019 to 2021 to identify hate speech in Hindi on various social media platforms. To categorize comments into abusive and non-abusive groups, several BERT models, including mBERT, DistilBERT, RoBERTa, HateBERT, and IndicBERT, have been utilized. Additionally, a comprehensive sentiment analysis of the derogatory comments has been performed. The research presents a stacked ensemble framework for binary (abusive and non-abusive) and multiclass (hate, offensive, and profane) classification that integrates predictions from mBERT, HateBERT, and IndicBERT models. Further, the study provides an integrated approach for providing abusive comment detection and sentiment analysis using a multi-output model. The proposed ensemble model achieves 94% accuracy in binary classification, with precision, recall, and F1 scores all approaching 94%. Multiclass ensemble models yield an accuracy of 93% and associated precision, recall, and F1 scores of 91%, 92%, and 92%, respectively. A comparative analysis using several state-of-the-art techniques has been generated to verify the efficacy of the suggested methodology.

Read the paper · More papers on PaperTik