Where Are All The Bugs?

Mary Holstege · Balisage series on markup technologies · 2013

In a large code and complex code base, it becomes unfeasible to manually develop tests for every feature and combination of features. The key to quality assurance in this context is automation and focus. Automatic generation of tests creates its own problems, however, as the execution of a complete cross-product of all interactions will take too long to execute, and small defects can give rise to large numbers of regression failures that must be manually analyzed. Manually identifying the interactions is itself a challenging undertaking, as is automatically generating meaningful test cases. It becomes important to make smart choices about what to expend effort on so as to minimize the risk of undetected code defects. This paper reports on an attempt to find areas to focus testing on in a large XQuery code base by performing XQuery introspection on that code base, treating the set of functions and parameter and return types as vocabularies, and computing TF-IDF scores over the terms in those vocabularies. To the extent that function names and types follow classic Zipf distributions, using TF-IDF scoring over those vocabularies makes mathematical sense. Terms that score high will be those that are common enough to be important (high term frequency) but not so ubiquitous that they tend not to be covered by other tests (high inverse document frequency).

Read the paper · More papers on PaperTik