Automatic Subspace Clustering: for High-Dimensional Data

Jiwu Zhao · Univ. Duesseldorf: Duesseldorfer Dokumenten- und Publikationsserver · 2014

Clustering is an important task of data mining. The purpose of clustering is discovering and grouping similar objects in a data set with the principle that objects in the same group (cluster) are similar. Meanwhile, the ones from different clusters are dissimilar. The traditional clustering approaches are designed for searching clusters in the entire space. However, in high-dimensional real world data sets, there are usually many irrelevant dimensions for clustering, where the traditional clustering methods work often improperly. Subspace clustering is an extension of traditional clustering that enables finding subspace clusters only in relevant dimensions within a data set. However, most subspace clustering methods usually suffer from the issue that their complicated parameter settings are almost troublesome to be determined, and therefore it can be difficult to implement these methods in practical applications. In this dissertation, we introduce two novel subspace clustering methods SUGRA (Subspace Clustering with the Gravitation Function) and ASCDD (Automatic Subspace Clustering with the Distance -Density Function). The first algorithm SUGRA takes a gravitation function to calculate the densities of objects. It searches clusters from low- to high-dimensional subspaces. The second algorithm ASCDD uses another density function and computes the density distribution directly in high-dimensional subspaces. The relevant subspaces are explored by comparing their entropy values. The clusters in ASCDD are searched with the technique of neighborhood expansion. Both of the subspace clustering methods are designed with the principle of uncomplicated parameter setting and easy applicability. For example, SUGRA can separate non-cluster objects by one threshold that is close to a constant. ASCDD requires only one simply determinable parameter in the step of the neighborhood expansion. Finally, we compare SUGRA and ASCDD with other subspace clustering methods in different empirical experiments with various aspects. The results show that the two proposed subspace clustering methods are accurate and easy applicable on different types of data sets.

Read the paper · More papers on PaperTik