Efficient detection of offensive social media comments in Assamese language using LSTM
Tulika Chutia, Sanjib Bora, Nomi Baruah, Debajani Baruah, Swarnangka Barman, Joy Anupol Neog, Bikokhita Dutta · 2025
The increasing incidence of abusive language on online sites has necessitated the formulation of automated detection systems to offset its damaging effects on individual users and social communication. This paper introduces an offensive sentence detection system for the Assamese language using LSTM networks. As no public dataset was available, a public dataset of 3,566 Assamese sentences was created and labelled as either offensive or non-offensive. The approach attained an accuracy of 88.10% and an F1 score of 87.80%, thus validating the potential of the mockup to analyse offensive content correctly. Comparative analysis with related research in other languages revealed that the proposed approach surpassed existing procedures, including SVM and XML-RoBERTa. Despite its success, the study admits that the size of the dataset is small and incapable of detecting subtle intentions for using offensive language. Upcoming work may incorporate increasing the size of the corpus and enhancing the ability of the model to capture subtle forms of offensive content. In general, this research shows that the area of NLP is possible for monitoring and filtering offensive comments in Assamese.