Advance Spam Detection Using Machine Learning

Pratyush Pratyush, Lovish Jain, Aklavya Aklavya, Ajeet Kumar Sharma · 2025

The rapid expansion of digital communication has resulted in a significant rise in unsolicited messages, commonly known as spam. This influx disrupts communication systems and consumes valuable resources. Although spam detection methods have advanced, ongoing challenges persist due to the constantly changing nature of spam and the unequal distribution of data, where legitimate (ham) messages far exceed spam in number. In this paper an optimize machine learning-based model is proposed for accurately classifying messages as either spam or ham. The model integrates the Synthetic Minority Over-sampling Technique (SMOTE) to handle data imbalance, ensuring more accurate and unbiased results to counteract dataset imbalance by creating synthetic samples of the minority class. This approach improves the model's ability to identify underrepresented spam messages effectively. The proposed model is more efficient and produces high accuracy of 98.38. Future research directions are also given to provide insights for researchers and security professionals.

Read the paper · More papers on PaperTik