Software Effort Estimation and Conclusion Stability
Tim Menzies, Omid Jalali, Jairus Hihn, D. R. Baker, Karen Lum · 2007
Abstract — This paper revisits the conclusion instability problem identified by Kitchenham, Foss, Myrtveit et.al.; i.e. conclusions regarding which software effort estimation method is “best ” is highly contingent on (1) the evaluation criteria and (2) the subset of the data used in the evaluation. Using non-parametric methods (the Mann-Whitney U test), we show how to avoid conclusion instability. This paper reports a study that ranked 158 effort estimation methods via three different evaluation criteria and hundreds of different randomly selected subsets. The same four methods were ranked higher than the other 154 methods regardless of which evaluation criteria or data subset was applied. Hence, we recommend non-parametric evaluation to evaluate and prune effort estimation methods. More specifically, when learning effort estimators from COCOMO-style data, we find that manual stratification defeats many complex algorithmic methods. However, we can do better than manual stratification by augmenting Boehm’s local calibration method with simple linear-time row and column pruning pre-processors. We also advise against model trees, linear regression, exponential time feature subset selection, and (unless the data is sparse) methods that average the estimates of nearest neighbors. To the best of our knowledge, this report is the first to offer stable conclusions regarding effort estimation across such a wide range of methods. Index Terms — COCOMO, effort estimation, data mining, evaluation, Mann-Whitney U test, non-parametric tests. I.