Quantitative Evaluation of Clustering Results Using Computational Negative Controls
Ronald K. Pearson, Tom Zylkin, James S. Schwaber, Gregory E. Gonye · 2004
Most partition-based cluster analysis methods (e.g., k-means) will partition any dataset D into k subsets, regardless of the inherent appropriateness of such a partitioning. This paper presents a family of permutation-based procedures to determine both the number of clusters k best supported by the available data and the weight of evidence in support of this clustering. These procedures use one of 37 cluster quality measures to assess the influence of structure-destroying random permutations applied to the original dataset. Results are presented for a collection of simulated datasets for which the correct cluster structure is known unambiguously.