A Measure of Description Quality for Data Mining and its Implementation in the AQ18 Learning System

Ryszard S. Michalski, Kenneth A. Kaufman · 1999

Given a sufficiently large database, it is usually possible to derive many different hypotheses about the data. Therefore, an important problem in data mining is which hypothesis to select, or, more specifically, how to define a criterion that best articulates the requirements of the given task domain. This paper presents an approach to this problem, in which a knowledge agent seeks descriptions that optimize a problem-oriented description quality measure. The proposed measure flexibly combines two criteria characterizing descriptions, completeness and consistency gain, into a single numerical measure of description quality Q(w). It is shown that by modifying the parameter w in Q(w), the proposed description quality specializes to different measures described in the literature. The Q(w) measure has been implemented in the AQ18 learning system, and compared to several other methods, such as Information Gain, PROMISE, CN2, IREP and RIPPER. A general measure of description utility can be obtained by integrating the description quality with description simplicity through a Lexicographic Evaluation Functional (LEF). Experimental results have demonstrated the generality and the flexibility of the proposed method. 1 The Problem Statement A typical objective in extracting knowledge from data is to hypothesize general rules or patterns that can be used effectively for predicting or classifying future data. Given a sufficiently large database, one can generate a large number of hypotheses characterizing the data. A question then arises as to what criteria will lead to a hypothesis with the maximum predictive accuracy, as well as other desirable properties, such as generality and simplicity. If one can make the assumption (usually unrealistic) that data contains no noise, the preconditions for admissibility

Read the paper · More papers on PaperTik