Principles for generalised testing of knowledge bases
Timothy James Menzies · UNSWorks (University of New South Wales, Sydney, Australia) · 2022
Modern KA views KBS construction as the construction of inaccurate surrogates models of reality. We argue that such potentially inaccurate models must be tested, lest they generate inappropriate output for certain circumstances. Testing can only demonstrate the presence of bugs (never their absence) and so must be repeated whenever new data is available. That is, testing is an essential, on-going process through-out the lifetime of a knowledge base. This view motivated our development of a general computational model for automatically testing models in vague domains. A vague domain is (i) poorly-measured; and/or (ii) lacks a definitive oracle; and/or (iii) is indeterminate/non-monotonic. Testing in such vague domains necessitates making assumptions about unmeasured variables and maintaining mutually exclusive assumptions in separate worlds. Most domains tackled by KBS are vague; i.e. our definition of test is widely applicable. Surprisingly, our generalised test engine is also a generalised inference engine. Domain-specific modeling constructs are mapped into a directed and-or graphs of edges E and vertices V. Consistent subsets of E are extracted which explain known observations in terms of known inputs to the model. This extraction process directly operationalises the model extraction process which Clancey and Breuker argue is the core of expert systems inference. We find that the second generation knowledge level modelling approaches (e.g. KADS) complicates and separates processes that can be unified and simplified via our abductive framework. Consequently, we propose replacing methodologies like KADS with our abductive architecture that supports and simplifies both inference and testing. Limits to model testing are also limits to model construction since potentially inaccurate models that can't be tested should not be used. We find that our process is practical for model at least up to $\vert V\vert$ = 850 and $\vert E\vert/\vert V\vert < 7.$ However, we lose the ability to test models in vague domains above $\vert E\vert/\vert V\vert$ = 7. Based on surveys of known fielded expert systems, we conclude that: (i) our technique is practical for the models seen in current KA practice; however, (ii) modern KA is teetering on the edge of model testability/construction.