Global Ordering For Multi-dimensional Data: Comparison with K-means Clustering
Baiyang Liu, Casimir A. Kulikowski · 2009
This paper describes a novel approach to estimate the quality of clustering based on finding a linear ordering for multi-dimensional data by which the clusters of the data fall into intervals on the ordering scale. This permits assessing the result of such local clustering methods like K-means so as to filter inhomogeneous or outlier clusters that can be produced. Preliminary results reported here indicate that the method is valuable to determine, in two dimensions, the number of visually perceived clusters generated by a mixture of Gaussian distribution model, corresponding to the number of actual generating distributions when the means are far apart, but corresponding to the reduced number of clusters arising from the perceived This paper presents a new type of ordering for multi-dimensional data which can help eliminate outliers for clustering results such as those generated by K-means procedures. It is based on the notion that goodness of clustering is related to finding similar solutions resulting from very different methods. Consistency of results from different methods is considered