The Impact of Sample Size after Sampling on the Accuracy of Machine Learning Models
Siyu Chen, Jinwen Zheng, Jinying Li · 2024
Sampling techniques are a commonly used method in the field of statistics and machine learning for selecting a portion of a sample from a totality for analysis. The main purpose of sampling techniques in this paper is to deal with imbalanced data in the classification models. In machine learning when the target variable, i.e., positive and negative samples are unbalanced, direct modeling will result in poor model learning ability on a few classes of samples, so sampling is usually required. Based on this research status of existing tools for machine learning such as under-sampling, over-sampling, mixed sampling, etc., this paper will study from three perspectives. First, the effect of different sample sizes through sampling on the model accuracy of the following modeling is explored, and it is found that increasing the sample size may improve the prediction effect of the model. Second, by changing the proportion of positive and negative samples to investigate the changes on the correctness of the model, and found that different positive and negative sample proportions will have an impact on the prediction effect of the model, increasing the proportion of positive samples may improve the accuracy of the model, and balancing the proportion of positive and negative samples may be more conducive to the performance of the model. Thirdly, by exploring the effects of different random seeds values on the model, it is concluded that different random number seeds can make impacts on the stability and generalization ability of the model. The research on this topic can effectively fill the gap in the field of machine learning sampling, and provide some inspirations for determining sample size after sampling as well as the proportion of positive and negative samples.