SMS spam detection using N-grams and bag of words techniques

V. Nagasree, CH Sai Chandrakanth, Mohammed Riyaz, R. Sai Chandra Ktran · 2025

Nowadays, spam messages have become alarming issues in this digital communication world. This paper discusses a comprehensive study of SMS spam detection using NLP techniques, specifically n-grams and the bag-of-words (BOW) model. We use a labeled dataset of SMS messages with the class label as spam or ham, representing non-spam. A model built from this data set using a predictive algorithm should clearly classify spam messages. Our process is divided into a number of stages. These include data collection, preprocessing, feature extraction, model training, and evaluation. The first steps are to clean the data we gather, first cleansing any raw text with standard pre-processing-lower casing, tokenization, and stop word removal. We employ the bag-of-words approach enhanced with n-grams, such that the model may capture single words as well as word sequences. This extends the feature space for consideration of the contextual information extracted. The resulting features are then used to train a Naive Bayes classifier, SVM, and Random Forest which is a relatively good algorithm for this task type because of its simplicity but strong performance in the classification task. The accuracy, precision, recall, and Fl score are used to assess the performance of the model. Our results depict the efficiency of using n-grams integration in order to advance the classification accuracy for the spam message. The work indicates the importance of more advanced approaches in text processing for combating SMS spam and provides a robust framework that may be further developed and adjusted according to applications made for text classification. Future work will explore further integration with more advanced NLP models and additional feature engineering to make the accuracy of detection even more improvement.

Read the paper · More papers on PaperTik