Detecting spam and ham SMS messages using natural language processing and machine learning algorithms
Hussein Abdel-Jaber · PeerJ Computer Science · 2025
Many people use mobile devices for communication, including the short message service (SMS). Attackers can exploit SMS to send spam messages and carry out phishing or malware attacks. These attacks pose risks to users by potentially damaging devices or stealing, altering, or deleting data. The purpose of this study is to detect spam and ham SMSs efficiently. The methodology of this research is to propose a data-driven process model based on natural language processing (NLP) and machine learning algorithms to detect spam and ham SMSs. The machine learning algorithms used in this model are K-nearest neighbors (KNN), decision tree (DT), random forest (RF), gradient boosting (GB), multi-layer perceptron (MLP), and support vector machine (SVM). The proposed model uses several steps, such as deleting punctuation from the dataset, converting data in the dataset into lowercase letters, applying tokenization on the data in the dataset, deleting English stopwords from the dataset, utilizing stemming or lemmatization, using the map function to replace class values with numeric values, using a feature extraction method (CountVectorizer), applying stratified shuffle split cross-validation, using the random oversampling method to solve the imbalanced classes in the dataset, and training and testing the models using the above-mentioned machine learning algorithms. The classification measures used for the considered machine learning algorithms are accuracy, precision, recall, and F1-score. The proposed machine learning models are compared using the classification measures based on balanced or imbalanced dataset to specify which model provides better classification performance. In addition, the proposed machine learning models are compared using the precision-recall curve and the area under the precision-recall curve based on an imbalanced dataset to evaluate the performance of the classification models of the compared algorithms. Moreover, the proposed model is compared with related works to assess its performance against previous works. The experimental results showed that multi-layer perceptron (MLP) offers the highest accuracy, precision, and F1-score results. In addition, KNN, RF, MLP, and SVM provide similar and the highest recall results. Based on the results of the areas under the precision-recall curves, the classification models of the compared algorithms are performing well.