Automated Learning of Decision Text Categorization
Chidanand V. Apte, Fred J. Damerau, Sholom M. Weiss · 1994
We describe the results of extensive experiments using optimized rule-based induction methods on large document collections. The goal of these methods is to discover automatically classifica-tion patterns that can be used for general document categorization or personalized filtering of free text. Previous reports indicate that human-engineered rule-based systems, requiring many man-years of developmental efforts, have been successfully built to “read ” documents and assign topics to them. We show that machine-generated decision rules appear comparable to human performance, while using the identical rule-based representation. In comparison with other machine-learning techniques, results on a key benchmark from the Reuters collection show a large gain in performance, from a previously reported 67 % recall/precision breakeven point to 80.5%. In the context of a very high-dimensional feature space, several methodological alterna-tives are examined, including universal versus local dictionaries, and binary versus frequency-related features.