Improving SMS Spam Detection Through Machine Learning: An Investigation of Feature Extraction and Model Selection Techniques

William Siagian, Melisa Rachel Setiadi, Simeon Yuda Prasetyo · 2023

When it comes to SMS, the topic of spam is a big obstacle that requires urgent attention. The majority of SMS messages that fall under the category of spam are commercial communications, which are obviously illegal and unethical, followed by phishing scams and messages that may disturb the user's peace of mind. Users may find this to be extremely inconvenient as it consumes energy to erase those messages, slows down websites, and may even infect computers with viruses. We must therefore distinguish between SMS messages that are spam and those that are not in order to prevent this. We are recognizing this in this project utilizing machine learning algorithms. In this paper, we will talk about the algorithms we employ, compare them all to the dataset we use, and finally select the algorithm that will be most effective at determining which SMS messages fall under the spam category. Based on comparisons of precision and accuracy, we choose the optimum algorithm, then for the results of the research in our paper, we found that the best non-pretrained machine learning model was GRU with the best feature extractor being GloVe, with an accuracy of 91.9171% and precision of 91.9187%. However, the pretrained machine learning model BERT still outperformed the non-pretrained model with an accuracy of 99.0166% and precision of 99.0179%.

Read the paper · More papers on PaperTik