Ensemble methods for the prediction of number of faults: A study on eclipse project

Santosh Singh Rathore, Sandeep Kumar · 2016

Software fault prediction using different machine learning and statistical techniques has been reported by various researchers. However, different techniques produced different results for software fault prediction and thus showed the performance bottleneck of single techniques. Moreover, most of the researchers focused on predicting software modules being faulty and non-faulty, i.e., binary class prediction of faults. On the other hand, in recent year, some researchers showed that ensemble methods produced the improved performance for software fault prediction compared to single fault prediction techniques. Motivated by this reason, we perform an empirical study of different homogeneous ensemble methods for the prediction of number of faults. The study includes bagging, boosting, random subspace, rotation forest, and stacking ensemble methods and uses three different techniques, linear regression, multilayer perceptron, and decision tree regression as the base learners for the ensemble. The experiments are performed for three different fault datasets corresponding to the Eclipse project. To our knowledge, very few works on the prediction of number of faults using Eclipse dataset have been reported. Results indicated that overall ensemble methods produced better performance than using a single fault prediction technique. Out of five ensemble methods, random subspace outperformed other ensemble methods. Rotation forest, bagging, and boosting performed moderately. Stacking performed relatively poor compared to other ensemble methods.

Read the paper · More papers on PaperTik