Data mining for text categorization with semi-supervised agglomerative hierarchical clustering

Antonio F. Skarmeta, Amine M. Bensaid, Nadia Tazi · International Journal of Intelligent Systems · 2000

In this paper we study the use of a semi-supervised agglomerative hierarchical clustering (ssAHC) algorithm to text categorization, which consists of assigning text documents to predefined categories. ssAHC is (i) a clustering algorithm that (ii) uses a finite design set of labeled data to (iii) help agglomerative hierarchical clustering (AHC) algorithms partition a finite set of unlabeled data and then (iv) terminates without the capability to label other objects. We first describe the text representation method we use in this work; we then present a feature selection method that is used to reduce the dimensionality of the feature space. Finally, we apply the ssAHC algorithm to the Reuters database of documents and show that its performance is superior to the Bayes classifier and to the Expectation-Maximization algorithm combined with Bayes classifier. We showed also that ssAHC helps AHC techniques to improve their performance. © 2000 John Wiley & Sons, Inc.

Read the paper · More papers on PaperTik