Random Testing
Richard G. Hamlet · Encyclopedia of Software Engineering · 2002
Abstract In computer science, originally in its rarefied offshoots centered around artificial‐intelligence laboratories at Stanford and MIT, the adjective “random” is slang with a number of derogatory connotations ranging from “assorted, various” to “not well organized” or “gratuitously wrong”. “Random testing” of a program in this sense describes testing badly or hastily done, the opposite ofsystematictesting such as functional testing or structural testing. This slang colloquial meaning, which might be rendered “haphazard testing” in normal parlance, is probably the common one, as in the sentence “Random testing, of course, is the most used and least useful method.” In contrast, the technical, mathematical meaning of “random testing” refers to an explicit lack of “system” in the choice of test data, so that there is no correlation among different tests. If the technical meaning contrasts “random” with “systematic,” it is in the sense that fluctuations in physical measurements are random (unpredictable or chaotic) vs. systematic (causal or lawful). It desirable to be “unsystematic” on purpose in selecting test data for a program. Because there are efficient methods of selecting random points algorithmically, by computing pseudorandom numbers; thus a vast number of tests can be easily defined. Also statistical independence among test points allows statistical prediction of significance in the observed results. In this article it will be seen that point (1) may be compromised because the required result of an easily generated test is not so easy to generate and (2) is the more important quality of random testing, both in practice and for the theory of software testing. To make an analogy with the case of physical measurement, it is only random fluctuations that can be “averaged out” to yield an improved measurement over many trials; systematic fluctuations might in principle be eliminated, but if their cause (or even their existence) is unknown, they forever invalidate the measurement. The analogy is better than it seems; in program testing, with systematic methods, we know what we are doing, but not what it means; only by giving up all systematization can the significance of testing be known.