Examining the Efficiency of Various Classification Algorithms in Spam Detection Datasets
Mohd. Umar, Mahima Gupta, Rishabh Didwania, Rajat Verma, Namrata Dhanda · 2024
This study investigates the relationship between the accuracy of Machine Learning (ML) models used to differentiate between Spam and Normal Messages for different dataset sizes. Efficiency evaluation of four classification models (Decision Tree, Passive Aggressive, Random Forest, and Regression) using 1500, 3000, and 5500 samples from three distinct dataset sizes is performed in this paper. The primary objective is to determine how variations in dataset size impact these models’ accuracy. This work not only quantifies performance shifts but also offers insights into how responsive these models are to changes in data amount through careful experimentation and analysis. These findings enhance understanding of ML model behaviour in identifying unwanted messages and provide practical guidance for selecting and configuring models based on available dataset sizes. They also help to improve the design of trustworthy and accurate categorization systems.