Hate Speech Detection for a low level language (Assamese)

Nomi Baruah, Arjun Gogoi, Rituraj Phukan, Pritom Jyoti Goutom · 2023

Numerous social media issues like "Cyber conflict" have evolved as a result of the online community’s explosive expansion, and they can negatively impact both individual and group social interactions. Social media networks struggle to control all of their users, necessitating the necessity for an automated hate-speech classifier. We tried to categorize hate words gathered from social media platforms like Facebook, YouTube, and Twitter despite a number of obstacles, including the absence of Unicode, text corpora, and other NLP tools in the Assamese language. The algorithms used in this work to examine hate speech in Assamese are Random Forest and Linear SVC. Additionally, we examine current works that address hate speech and antisocial behavior in all Indian languages. Our results show that both algorithms perform well, with Random Forest surpassing Linear SVC by 3.33 percent and an accuracy rate of 88.33 percent. We compared our suggested techniques to current studies in other Indian languages using both lemmatized and non-lemmatized versions of words, as there has been no study done in the Assamese language for Hate-Speech identification.

Read the paper · More papers on PaperTik