Gender Abusive Bengali Text Classification using Enhanced CNN-LSTM Model with Bangla BERT Base Preprocessing
Mayeesha Farjana, Shyla Afroge, Azmain Yakin Srizon · 2024
With the rapid proliferation of social media and online communication platforms, the presence of gender-based abusive language has emerged as a significant issue in the digital sphere. Addressing this challenge necessitates the development of effective Natural Language Processing (NLP) techniques capable of accurately detecting and categorizing abusive content to safeguard users from harmful interactions. This study introduces a novel approach for Gender-Based Abusive Bengali Text Classification, which plays a crucial role in fostering safer online environments. The proposed methodology employs an enhanced CNN-LSTM model in conjunction with Bangla BERT Base for preprocessing, resulting in an impressive accuracy of 97.94%, outperforming prior models and demonstrating its effectiveness in identifying gender-targeted abusive language within the Bengali-speaking community. The primary contribution of this work lies in the development of a robust and precise classifier for gender-based abusive Bengali text, which holds significant potential for mitigating harmful online behavior. Consequently, this research not only advances the domain of NLP but also supports the creation of a safer and more inclusive digital space for Bengali language users.