Application of BIRCH to text clustering

Ilia Karpov, Alexandr Goroslavskiy · 2012

This work represents a clustering technique, based on the Balanced Iterative Reducing and Clustering using Hierarchies (BIRCH) algorithm and LSA-methods for clustering large, high dimensional datasets. We present a document model and a clustering tool for processing texts in Russian and English languages and compare our results with other clustering techniques. Experimental results for clustering the datasets of 10’000, 100’000 and 850’000 documents are provided.

Read the paper · More papers on PaperTik