Optimizing SMS Spam Detection with Large Language Models and Transformer Architectures
Mohamed H. Ahmed, A. S. M. Shakil Ahamed, Fahim Shakil Tamim · 2025
In recent years, the Short Message Service (SMS) has become omnipresent, with most people ignoring emails while nearly all check their daily text messages. These messages may contain spam, providing useless promotions or links that could lead to the installation of malicious applications into the system. Spam detection involves identifying minute linguistic signs and patterns in text messages, with detecting disguised spam presenting a significant challenge due to the continuous evolution of techniques. It aims to develop an intelligent system for classifying messages into two different classes, namely Spam or Ham from English SMS spam messages. In order to do so, this study explores different fine-tuned machine learning (Logistic Regression-LR, Multinomial Naive Bayes-MNB, Support Vector Machine-SVM) models, deep learn- ing(Convolutional Neural Network-CNN, Bidirectional Long Short-Term Memory-BiLSTM, CNN+BiLSTM) models, transformer(M-BERT, BERTbase, XLM-Rbase, XLNetbase) techniques and two large language models(Phi-3 and H2O-Danube). Finally, we experimented with two large language models, Phi-3 and H2O-Danube. Out of all the models we tested, H2O-Danube outperformed all other models with a macro F1-score of 0.94, proving to be the best for SMS spam detection, surpassing traditional machine learning, transformer, and LLM models.