Sample Size on the Impact of Imbalance Learning
Wei Mei Zhi, Hua Ping Guo, Ming Fan · Advanced materials research · 2013
Classification of imbalanced data sets is widely used in many real life applications. Most state-of-the-art classification methods which assume the data sets are relatively balanced lose their efficiency. The paper discusses the factors which influence the modeling of a capable classifier in identifying rare events, especially for the factor of sample size. Carefully designed experiments using Rotation Forest as base classifier, carried on 3 datasets from UCI Machine Learning Repository based on weak show that, in particular imbalance ratio, increases the size of training set by unsupervised resample the large error rate caused by the imbalanced class distribution decreases. The common classification algorithm can reach good effect.