A Comparative Study: Using Machine Learning and Transformers Model to Identify Spam in SMS
Lour Atwe, Mahmmoud Jazzar, Amna Eleyan, Tarek Bejaoui · 2025
In this modern era, the identification of SMS spam is crucial due to the significant risk that spam poses to users. In this research, various supervised machine learning models (Support Vector Machine (SVM), Naïve Bayes, and Random Forest) and transformer-based models (RoBERTa, DistilBERT) were utilized and trained on a relatively large, new dataset (super_sms_dataset) that was published in early 2024, which reached 67k records. In order to compare the performance of these models on different datasets, the same models were run on a different, smaller, and common dataset, the UCI dataset, which contains 5574 records. Consequently, transformer-based models outperformed traditional machine learning models, with the RoBERTa model achieving an impressive performance of 99.46% accuracy on the “super_sms_dataset” dataset.