Characteristic concept representations

Piew Datta · 1997

In machine learning, concepts have traditionally been represented and learned using algorithms that represent only those characteristics that discriminate between two or more classes. Representations such as decision trees and rules provide increased classification ability, but do not provide a general characteristic description of the concepts. This dissertation explores the use of characteristic descriptions for concepts in the classification task. In our research, characteristic descriptions are represented by prototypes. The methods described focus on the issue of learning multiple prototypes for a class when necessary. Two diverse methods are developed. The first method, PL (Prototype Learner), attempts to separate examples into smaller subgroups with similar characteristics. The second method, SNMC (Symbolic Nearest Mean with Clustering), applies clustering techniques to separate groups of examples, using classification accuracy on the training set as its heuristic. The groups of examples in each of these algorithms are simplified into prototypes representing the examples and a nearest neighbor classification approach is used for class prediction. Our empirical results show that SNMC has an increase in average classification accuracy of about 2% over C4.5 and PEBLS on 20 domains from the UCI data repository. The experimental results on the UCI domains showed that SNMC classifies the best in average rank and accuracy of the six algorithms compared. PL classifies favorably to C4.5 and PEBLS on the same domains. These results show that prototype concept representations can be successfully applied to the classification task. The last portion of this dissertation introduces a task, inductive inference, which is a generalization of classification. In this task the algorithm must make predictions about any of the attribute values. We discuss five subtasks of the inductive inference task, including the classification task. We also discuss metrics that can be used to evaluate algorithms on these tasks. We contend that these metrics used in conjunction with classification accuracy provides a more general evaluation methodology than classification accuracy alone. We provide comparisons among PL, C4.5, PEBLS, and COBWEB.

Read the paper · More papers on PaperTik