Text Classification Using WordNet Hypernyms

Sam Scott, Stan Matwin · 1998

This paper describes experiments in Machine Learning for text classification using a new representation of text based on WordNet hypernyms. Six binary classification tasks of varying difficulty are defined, and the Ripper system is used to produce discrimination rules for each task using the new hypernym density representation. Rules are also produced with the commonly used bag-of-words representation, incorporating no knowledge from WordNet. Experiments show that for some of the more difficult tasks the hypernym density representation leads to significantly more accurate and more comprehensible rules. 1. Introduction The task of Supervised Machine Learning can be stated as follows: given a set of classification labels C, and set of training examples E, each of which has been assigned one of the class labels from C, the system must use E to form a hypothesis that can be used to predict the class labels of previously unseen examples of the same type [Mitchell 97]. In machine learnin...

Read the paper · More papers on PaperTik