Mining Clustering Dimensions

Sajib Dasgupta, Vincent Ng · 2010

Many real-world datasets can be clustered along multiple dimensions. For example, text documents can be clustered not only by topic, but also by the author’s gender or senti-ment. Unfortunately, traditional clustering algorithms produce only a single clustering of a dataset, effectively providing a user with just a single view of the data. In this paper, we propose a new clustering algorithm that can discover in an unsupervised manner each clustering dimension along which a dataset can be meaningfully clustered. Its ability to reveal the important clustering dimensions of a dataset in an unsupervised manner is par-ticularly appealing for those users who have no idea of how a dataset can possibly be clus-tered. We demonstrate its viability on several challenging text classification tasks. 1.

Read the paper · More papers on PaperTik