An Under-sampling Method: Based on Principal Component Analysis and Comprehensive Evaluation Model
Fu Yangzhen, Hong Zhang, Yaxin Bai, Weixuan Sun · 2016
Machine learning method can be used to forecast software fault, and identify modules which have the tendency to cause soft-error at the early life cycle, then software developer can modify these defect modules early. It has an important significance on the improvement of software reliability. However, fault samples of software data sets are smaller in number (one or two order of magnitudes) compared with fault-free samples, learning machine's predictive ability to fault samples has been restrained by this kind of unbalanced data sets. This paper put forward an under-sampling method based on principal component analysis (PCA) and comprehensive evaluation model to get rid of redundant majority class samples under the premise of conserving data of majority class characteristic as far as possible, and reaches to a balance between this two kinds of samples.