Using simulated data sets to compare data analysis techniques used for software cost modelling

Lesley M. Pickard, Barbara Ann Kitchenham, Steve Linkman · IEE Proceedings - Software · 2001

The goals of this study were to compare different data analysis methods and to demonstrate the viability of simulation as a mechanism to allow such comparisons. Simulation was used to create data sets with a known underlying model and with non-Normal characteristics that are frequently found in software data sets: skewness, unstable variance, and outliers and combinations of these characteristics. Three data analysis approaches were investigated: residual analysis; multiple regression; classification and regression trees (CART). In addition to the standard statistical `least squares' version of each method, robust and non-parametric versions of the techniques were also investigated.It was found that standard multiple regression techniques were best if the data only exhibited moderate non-Normality. As might be expected, under more extreme conditions such as severe heteroscedasticity, the non-parametric techniques performed best. It was more surprising to find that under strongly non-Normal conditions the robust and non-parametric residual analysis techniques performed as well as the conventional robust and non-parametric versions of multiple regression.However, the most important result of this study is to demonstrate the value of simulation as a technique for evaluating different data analysis techniques under controlled conditions.

Read the paper · More papers on PaperTik