A Quality Metric for K-Means Clustering

M. Thulasidas · 2018

From a teaching perspective, K-Means algorithm for clustering figures in the introductory courses in data analytics because of its conceptual simplicity. However, it suffers from a couple of drawbacks in terms of variable selection and the determination of the optimal number of clusters. In this paper, we present a new, mathematically defensible, quality metric for K-Means clustering based on the standard score of the distribution of the centroids. Furthermore, we demonstrate how this Standard Score Metric (SSM) can be used for automatic variable selection and optimal number of clusters using well-known data sets as well as real data collected locally.

Read the paper · More papers on PaperTik