The Study of Model Generalization Ability for Spam Classification Based on Machine Learning Models
Shihao Zhang · 2024
Due to the widespread use of machine learning models in various third-party libraries in languages, such as the sklearn library, people can quickly learn how to simply use machine learning models to empower the projects. However, in many times, most people simply take the model and apply it, never thinking about its generalization ability in the projects where it belongs to, that is, whether the performance of the model will continue to be excellent in different application scenarios and completely different dataset structures. This work mainly tested the generalization ability of two models. By using a dataset for model training, a brand-new dataset is introduced into the trained model to test its generalization ability. If it is found that the generalization ability is not high, our model is optimized to improve its generalization ability as much as possible. After optimization, their generalization ability (represented by precision improvement here, as explained below) has significantly improved, from 38.56% to 100.00% and from 40.24% to 54.79% respectively.