Simultaneous categorization of text documents and identification of cluster-dependent keywords
Hichem Frigui, Olfa Nasraoui · 2003
We propose an approach to clustering text documents based on a coupled process of clustering and cluster-dependent keyword weighting. The proposed approach is based on the the fuzzy c-means clustering algorithm. Hence it is computationally and implementationally simple. Moreover, it learns a different set of keyword weights for each cluster. This means that, as a by-product of the clustering process, each document cluster will be characterized by a possibly different set of keywords. The cluster dependent keyword weights help in partitioning the document collection into more meaningful categories. They can also be used to automatically generate a brief summary of each cluster in terms of not only the attribute values, but also their relevance. For the case of text data, this approach can be used to automatically annotate the documents. We illustrate the performance of the proposed algorithm by using it to cluster a real collection of text documents.