A Dimensionality reduced Text data clustering with prediction of optimal number of clusters

M. Ramakrishna Murty, J. V. R. Murthy, Prasad Reddy, Suresh Chandra Satapathy · International Journal of Applied Research on Information Technology and Computing · 2011

Grouping large number of text documents is a challenging task due to high dimensional representation of the vector space model. Higher dimensionality of text data leads to computational burden and inefficient cluster results. In this work we improve the quality of the text document clustering using Singular Value Decomposition technique along with dimensionality reduction. In this paper we have also proposed Singular Value Decomposition which helps in finding the appropriate number of clusters for k-means clustering technique according to the singular values (Eigen values) calculated in the Singular Value Decomposition method.

Read the paper · More papers on PaperTik