A general framework for clustering high-dimensional datasets

Yanchang Zhao, Song Junde · 2004

In many fields, the datasets used in data mining applications are usually of high dimensionality. Most existing algorithms of clustering are effective and efficient when the dimensionality is low, but their performance and effectiveness degrade when the data space is high-dimensional. One reason is that their complexity increases exponentially with the dimensionality. To solve the problem, we put forward a general framework for clustering high-dimensional datasets. Common clustering algorithms, when combined with our framework, can be applied to cluster high-dimensional datasets efficiently. In our framework, a high-dimensional clustering is broken into several one- or two-dimensional clustering phases. During each phase, only one or two dimensions are involved. In such a way, common algorithms for clustering low-dimensional datasets can be used to process high-dimensional ones. In addition, attributes of different types can be processed with different algorithms in separate phases and datasets of hybrid data types can be handled easily. The efficiency and effectiveness of our framework is proven in our experiments.

Read the paper · More papers on PaperTik