Spam Detection For Emails Using Natural Language Processing And Explainable Machine Learning
Sathvika Patha, Vemula Pranay, Shiva Ram Poola, Naveen Kumar Penjarla, Mahesh Reddy Kandala · International Journal of Research Publication and Reviews · 2025
Every form of electronic communication faces the curse of spam-whether it be phishing, online fraud, or malware dissemination.Here, an NLP-style spam filter system has been developed that combines some text preprocessing techniques with understandable machine learning interpretability to actually classify emails with accuracy and interpretable perspectives.In this modeling, text feature embedding is done using TF-IDF vectorization to weed out "irrelevant" information to the spam classification, and the Synthetic Minority Over-sampling Technique (SMOTE) is introduced to care for the class imbalance problem.The comparison of models for performances based on evaluation metrics of accuracy, precision, recall, F1 score, and AUC-ROC is done for algorithms SVM and RF.On the merits of experimental results, SVM was able to rank itself first against RF with an accuracy of 96% as compared to 94%, which is of utmost importance given the precision-recall trade-off.The shadows on the keywords that lead to spam classifications were shown through the SHAP method, thus affording greater explanations for the model and augmenting the supported predictions.Thereafter, the interface is created using Streamlit for real-time spam detection for an end-user.Linking abstract methods of pre-processing NLP with explainable AI, this work becomes a lightweight and efficient solution looking far better than deeper learning methods.The present approach is, however, far too scalable to enhance email security on the business and consumer fronts.