Transformer-Based Hate Speech Detection in Assamese
Koyel Ghosh, Debarshi Sonowal, Abhilash Basumatary, Bidisha Gogoi, Apurbalal Senapati · 2023
Hate speech or hate content detection in a text is one of the emerging research in Natural Language Processing. Over several years, the popularity of social media has sky-rocketed, where anyone can share their feelings, desire, anger, etc. Sometimes, these posts include offensive information that hurts others, intentionally or unconsciously. Depending on the forum's visibility, such posts instantly spread over millions of people, which may lead to violence. In these circumstances, it is an essential task to identify whether the comment is hateful. In the Indian language context, this task presents greater complexity due to its multilingual nature and limited availability of resources. This situation motivated us to analyze hate speech in the Assamese language. The paper's contribution is in two aspects: firstly, it involves the development of a tagged dataset in Assamese comprising 4,000 sentences; secondly, by fine-tuning the mBERT-cased (bert-base-multilingual-cased) and the Bangla-BERT (sagorsarker/bangla-bert-base) model using the Assamese data to detect instances of hate speech effectively. Based on what we know, it is one of the pioneering attempts at detecting hate speech in the Assamese language.