Average Case Coverage for Validation of AI Systems

Tim Menzies, Bojan Čukić · 2001

Test engineers are reluctant to certify AI systems that use nondeterministic search. The standard view in the testing field is that nondeterminism should be avoided at all costs. For example, the SE safety guru Nancy Leveson clearly states “Nondeterminism is the enemy of reliability” (Leveson 1995). This article rejects the pessimism of the test engineers. Contrary to the conventional view, it will be argued that nondeterministic search is a satisfactory method of proving properties. Specifically, the nondeterministic construction of proof trees exhibits certain emergent stable properties. These emergent properties will allow us to rely that, on average, nondeterministic search will adequately probe the reachable parts of a search space. Our argument assumes that properties are proved using randomized search from randomly selected inputs seeking a randomly selected goal (Menzies & Cukic 1999; 2000; Menzies et al. 2000; Menzies, Cukic, & Singh 2000; Menzies & Singh 2001; Menzies & Cukic 2001). Our analysis applies to any theory that can be reduced to the negationfree horn clauses of Figure 1, plus some nogood predicates that model incompatibilities. Before beginning, we pause for an important caveat. This paper presents an average-case analysis of the recommended effort associated with testing. By definition, such an average case analysis says little about extreme cases of high critically. Hence, our analysis must be used with care if applied to safety-critical software.

Read the paper · More papers on PaperTik