Implementation of Data Sampling in Class Imbalance Learning for Cross Project Defect Prediction: An Empirical Study
Lipika Goel, Mayank Sharma, Sunil Kumar Khatri, Divya Damodaran · 2018
Cross-project defect prediction is a way in which the prediction model is trained by using the data sources of the different projects and then it is tested on the target data source. The data source from the different projects generates to a highly imbalanced source dataset. There exists an imbalance between the defect prone and non defect prone classes. This in turn degrades the performance of the predictive model. This paper performs an empirical analysis in a two-fold manner. Firstly it evaluates whether data sampling using SMOTE algorithm can improve the performance of the predictive model for different categories of cross project defect prediction (CPDP). Secondly, it also ensures whether this technique applied in CPDP is comparable to within project defect prediction (WPDP). Ensemble learning classifiers i.e. Gradient Boosting is used as the predictive model over 16 publically available datasets. The experimental results infer that SMOTE algorithm can be applied to overcome the problem of class imbalance on different categories of CPDP. Besides this it gives comparable results to WPDP with statistical significance.