A Comparison of Manual and Automatic Constructions of Category Hierarchy for Classifying Large Corpora.
Fumiyo Fukumoto, Yoshimi Suzuki · 2004
We address the problem dealing with a large collection of data, and investigate the use of automatically constructing category hierarchy from a given set of categories to improve classification of large corpora. We use two wellknown techniques, partitioning clustering, k- means and a loss function to create category hierarchy. k-means is to cluster the given categories in a hierarchy. To select the proper number of k, we use a loss function which measures the degree of our disappointment in any differences between the true distribution over inputs and the learner's prediction. Once the optimal number of # is selected, for each cluster, the procedure is repeated. Our evaluation using the 1996 Reuters corpus which consists of 806,791 documents shows that automatically constructing hierarchy improves classification accuracy.