Harnessing Deep Learning for Assamese Sarcastic Comment Detection from Social Media Text
Tulika Chutia, Bikokhita Dutta, Nomi Baruah · Procedia Computer Science · 2025
This paper focuses on developing a robust sarcasm detection system for Assamese, a language with rich cultural and linguistic nuances. The total number of comments collected is 5,497, out of which 2,997 are non-sarcastic comments and 2,500 are sarcastic. The collected dataset is further preprocessed to remove all punctuation, tokenize the data, and apply word embedding to enhance text quality for model training. Here the two architectures, CNN+LSTM and LSTM+Bi-LSTM, are sarcasm comment classifiers. In the CNN+LSTM model, a 1D convolution layer is stacked to capture local patterns; and then the LSTM layers are handled for long-term dependencies. For the LSTM+Bi-LSTM model, it uses both LSTM and bidirectional LSTM layers to handle the sequence data as well. Both architectures are tested using accuracy, precision, recall, and F1-score.The performance of the CNN+LSTM model was superior with 67% across all metrics, while the LSTM+Bi-LSTM had achieved 65%. This study reveals the effectiveness of hybrid deep learning architectures not only in the detection of sarcasm but also marks a significant contribution toward processing the Assamese language.