Comparing software fault predictions of pure and zero-inflated Poisson regression models
Taghi M. Khoshgoftaar, Kehan Gao, Robert M. Szabo · International Journal of Systems Science · 2005
Predicting the software quality prior to system tests and operations has proven to be useful for achieving effective reliability improvements. Poisson (pure) regression modelling is the most commonly used count modelling technique for predicting the expected number of faults in software modules. It is best suited to when the distribution of the fault data (dependent variable) is not biased, that is equidispersed fault data, whose mean equals the variance. However, in software fault data we often observe a large portion of zeros (no faults), especially in high-assurance systems. In such cases a pure Poisson regression model (PRM) may yield inaccurate fault predictions. A zero-inflated Poisson (ZIP) model changes the mean structure of a PRM, resulting in improved predictive quality. To illustrate the same, we examined software data collected from a full-scale industrial software system. Fault prediction models were calibrated using both pure Poisson and ZIP regression techniques. To prevent claims based on a biased data split (for the fit and test data sets), the data set was randomly split 50 times, and models were calibrated using each of these split combinations. A comparative hypothesis test between the pure Poisson and ZIP modelling techniques was performed. The test revealed that the ZIP model fitted better than its counterpart. Our comprehensive empirical comparative study presented in this paper showed that the ZIP model yielded better predictions than the PRM and also demonstrated better robustness in prediction accuracy across the 50 data splits.