Clustering by means of unsupervised decision trees or hierarchical and K-means-like algorithm

Patrice Bellot, Marc El-Bèze · 2000

A classical information retrieval system returns a list of documents to a user query. The answer list is often so long that users cannot explore all the documents retrieved. A classification of the retrieved documents allows to thematically organize them and to improve precision. In this paper, we present and compare two text classification algorithms. The first one is a clustering algorithm (K-Means-like) initialized with a partial hierarchical classification. The second one is a new algorithm that relies on unsupervised decision trees (UDTs). The indexing methods we use (TF-IDF weighting scheme, cosine similarity in the vector space model) prevent from really considering all the subjects dealt with in the texts. A better way to take all the themes into account is to cluster sentences from documents instead of documents as a whole. This is achieved the second method we propose. The effectiveness of these methods is evaluated over Amaryllis'99 corpora and queries. Since these methods ...

Read the paper · More papers on PaperTik