A Comparative Study of Machine Learning and Deep Learning Approaches for Identifying Assamese Abusive Comments on Social Media

Tulika Chutia, Nomi Baruah, Paramananda Sonowal · Procedia Computer Science · 2025

This paper addresses the growing concern of abusive language on online digital media, specifically focusing on the Assamese language. A new dataset of 2,669 Assamese social media comments was developed, comprising 1,329 labeled as abusive and 1,339 as non-abusive. The study weighs up the efficacy of traditional machine learning (ML) models Forest, Naive Bayes classifier, Support Vector Machines (SVM), and a deep learning (DL) model using Long Short-Term Memory(LSTM) networks and Recurrent Neural Network (RNN). The dataset creation involved rigorous data preprocessing, including the elimination of punctuation marks, stop words eliminations, emojis, and the tokenization and embedding of words. The Machine Learning models achieved varying levels of accuracy, with Random Forest performing the best among them. However, the LSTM network significantly outperformed all traditional ML models, demonstrating superior precision (87.83%), recall (89.88%), F1-score (88.84%), and accuracy (89.13%), and the RNN achieved 83.72% Precision, 84.05% Recall, 83.88% F1-score, and 84.46% Accuracy. These findings highlight the effectiveness of LSTM in capturing complex patterns and context in Assamese text, making it a robust tool for detecting abusive language. The research contributes valuable resources and insights for enhancing online safety in Assamese-speaking communities.

Read the paper · More papers on PaperTik