SNCS: Experimental Procedures to Develop Learning Based SMS Spam Identification Model by Using Superficial Neural Classification Strategy

Anitha G. S., Chakravarthi Bachina, S.V. Achuta Rao, N.P.G. Bhavani, Bonde A, Khaled E. Al-Qawasmi · 2024

The increasing use of mobile communication has made SMS an essential medium for interaction, but it has also led to the rise of spam messages that pose significant security and privacy risks. This research proposes the development of an SMS spam identification model using a Superficial Neural Classification Strategy (SNCS), integrating both traditional machine learning (ML) and deep learning (DL) techniques. Datasets were preprocessed using tokenization, text cleaning, and word embedding, followed by feature extraction through Term Frequency-Inverse Document Frequency (TF-IDF). Traditional ML models such as Naive Bayes and Support Vector Machines (SVM) were implemented alongside DL models like Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU). The results demonstrated that DL models outperformed traditional models, with the CNN-LSTM achieving the highest accuracy at 95.12%, followed by LSTM at 94.55%. Traditional ML models also performed well, with Gradient Boosting achieving an accuracy of 90.12%. Furthermore, addressing the class imbalance in the dataset improved the model's performance significantly, enhancing recall and F1-scores across all models. This study highlights the importance of model selection and data balancing in improving spam detection accuracy.

Read the paper · More papers on PaperTik