Less is More: Non-Redundant Subspace Clustering
Ira Assent, Emmanuel Müller, Stephan Günnemann, Ralph Krieger, Thomas Seidl · 2010
Clustering is an important data mining task for grouping similar objects. In high dimensional data, however, effects attributed to the “curse of dimensionality”, render clustering in high dimensional data meaningless. Due to this, recent years have seen research on subspace clustering which searches for clusters in relevant subspace projections of high dimensional data. As the number of possible subspace projections is exponential in the number of dimensions, the number of possible subspace clusters can be overwhelming. In this position paper, we present our work on identifying non-redundant, relevant subspace clusters which reduce the result set to a manageable size. We discuss techniques for evaluating, visualizing and exploring subspace clusterings, and propose some directions for future work. 1.