An empirical study on optimization of training dataset in harmfulness prediction of code clone using ensemble feature selection model
Sheng Yan, Liping Zhang, Dongsheng Liu · 2018
In order to solve the problem of irrelevant features and imbalanced data classification in the process of clone code harmfulness prediction, an integrated classifier algorithm based on RUS (Random Under Sampling) and Wrapper was proposed. Firstly, the majority of samples in training dataset were re-sampled into several proportional minority class data set, which were combined with minority samples to create multiple different training sample subsets; Then, a sequential floating forward search algorithm based Wrapper was proposed to select optimal feature subsets; The different proportions of training subsets were mapped with the corresponding optimal feature subsets; Finally, random forest classifier was used to evaluate the acquired optimized training dataset. The experimental results showed that this integrated classifier algorithm applied to code clone harmfulness prediction increased average about 7% in accuracy, F1 measure and AUC evaluation index. And compared with four other similar optimization methods, the AUC value of integrated classifier algorithm was increased by 10.3%, which expressed the feasibility and effectiveness of the ensemble feature selection model.