Inductive learning in the presence of uncertainty
Keith C. C. Chan · 1989
Inductive learning is concerned with the derivation of general laws by examining particular instances. It can at least be distinguished into two types: classification and prediction. Classification is concerned with the learning of a general class description from instances of that class whereas prediction is concerned with the learning of a general description of an ordered sequence of objects based only on a fragment of it. Traditionally, inductive learning systems, be they for classification or prediction, were developed by AI researchers under the assumption that the learning environment is noise-free. These systems are, therefore, unable to handle real-world problems that are characterized by the presence of various noise sources. In order to deal with uncertainty in classification and prediction tasks, some existing techniques for probabilistic inference can be employed. Unfortunately, many of these techniques can only be used with continuous-valued or binary-valued data. To effectively handle symbolic information represented as discrete multivalued data, the Dependence-Tree method was developed. This method, however, suffers from the problem of being 'variable-directed' (i.e. it can only determine what variables, but not how they are, dependent on each other). For this reason, an improved probabilistic inference method, the Event-Covering technique, has been proposed. Even though the performance of this method has been shown to be quite satisfactory, some theoretical and conceptual deficiencies have been discovered. In this dissertation, I first discuss what these deficiencies are. Then, based on the idea of residual analysis in statistics and the weighing of evidence in information theory, we present a new probabilistic inference method. This method, which is known as the Probabilistic Inference Technique (PIT), is simple yet very effective in uncovering hidden patterns in noisy data and is, therefore, able to overcome some of the problems that the commonly-employed methods face. Based on PIT, two learning systems, APACS and OBSERVER-II have been developed to deal with the problem of uncertainty in classification and prediction tasks, respectively. The performance of these systems have been evaluated using both simulated and real-world data. Compared to other existing systems, APACS is able to perform better both in terms of classification accuracy and computational efficiency whereas OBSERVER-II is, to the best of my knowledge, the only learning system that can handle the prediction problem in the presence of uncertainty.