Evaluating Text Mining and Classification Algorithms for Spam Detection

E. Loganathan, T. Mathankumar, P. Kalyanasundaram, P. Anitha, R. Geetha, J.K. Kanimozhi · 2025

Aim: 'The purpose of this study is to study and compare the performance of the most common methods of text mining and classification algorithms for spam detection, especially the comparison of the Naive Bayes classifier and the Decision Tree classifier.’ Materials and Methods: Two methods were evaluated for this analysis: The Decision Tree algorithm was set as an intervention for Group 1. These Decision Tree classes produce a model in the form of a wooden structure, where each node represents a function (word), and branches represent the decision rules that take the final classification of an email, such as spam or non-spam. This method gained accuracy from 80.0% to 89.0% and offers a more explanatory model that captures complex decision rules to classify e-post. Group 2 used Naive Bayes classification as a comparative method. This potential model uses the theorem of Bayes to classify the e-post based on conditional opportunities and word frequency. With an accuracy of 88.0% to 95.0%, this method simplifies calculations and allows for effective treatment of large data sets to assume that properties (words) are independent of each other. Results: The Naive Bayes model achieved 93% accuracy, which is much higher than the 81% of the decision tree model. It also showed better performance in different spam patterns and rapid classification time. Conclusions: Naive Bayes improves the decision tree in spam detection, with a high accuracy range (88.0% to 95.0% to 80.0% to 89.0%). It provides better accuracy, precision, recall, and F1 score and lowers false positives and negatives effectively. The statistical significance of this improvement is supported by a p-value < 0.05, indicating that the observed differences in performance are not due to chance but rather a meaningful enhancement in classification accuracy.

Read the paper · More papers on PaperTik