Spam No More: A Cross-Model Analysis of Machine Learning Techniques and Large Language Model Efficacies
Robin Chataut, Aadesh Upadhyay, Yusuf Usman, Mary Nankya, Prashnna Gyawali · 2024
With the increasing sophistication of phishing scams, financial fraud, and malicious cyber-attacks, the need for effective spam detection mechanisms to safeguard users is more critical than ever. In this paper, we present a comprehensive evaluation of traditional machine learning models and Large Language Models (LLMs) in the context of spam detection. By assessing a variety of traditional ML models such as Support Vector Machines (SVM), Logistic Regression, Random Forest, Naive Bayes, K-Nearest Neighbors (KNN), and XGBoost on several performance metrics, we establish a baseline of effectiveness for spam identification tasks. We extend our analysis to include LLMs, specifically ChatGPT 3.5, Perplexity AI, and our own customized fine-tuned GPT model, referred to as TextGPT. Our findings show that while traditional ML models are effective, LLMs demonstrate exceptional potential in enhancing spam detection. Through a rigorous comparative analysis, this study highlights the strengths of both traditional and advanced approaches, showcasing the promising application of LLMs in improving spam detection processes.