Assessing Predictors of Software Defects
Tim Menzies, Justin Distefano, Andres Orrego · 2004
OVERVIEW: When learning defect detectors from static code measures, NaiveBayes learners are better than entrophy-based decision-tree learners. Also, accuracy is not a useful way to assess those detectors. Further, those learners need no more than 200-300 examples to learn adequate detectors, especially when the data has been heavily stratified; i.e. divided up into sub-sub-sub systems (and by “adequate”, we mean that those detectors perform nearly as well slower, more expensive manual inspections). The rest of this paper describes how we reached those conclusions after (i) an analysis of known baselines in the literature and (ii) a review of a assessment methods for detectors learned from the NASA defect logs of Figure 1. BASELINES: If defect detectors are interesting, they must somehow be better than known baselines in the literature. For example, consider manual code reviews. These reviews are labor intensive; depending on the methods, 8 to 20 LOC/minute can be inspected and this effort repeats for all members of the review team, which can be as large as four or six [5]. These reviews can also be effective. A recent panel at IEEE Metrics 2002 [6] concluded that such reviews can find ≈60 % of defects1. That defect detection rate has a wide variance. Raffo found that the defect detection capability of industrial inspection methods can vary from T R(35, 50, 65) % 2 for full Fagan inspections, to T R(13, 21, 30) for some widely-used industrial practices. PUBLIC DOMAIN PROBLEMS: The data used in this study comes from the CM1, JM1, PC1, KC1 and KC2