K-MEANS CLUSTER ANALYSIS AND MAHALANOBIS METRICS: A PROBLEMATIC MATCH OR AN OVERLOOKED OPPORTUNITY?
Andrea Cerioli, Dipartimento di Economia · 2005
In this paper we consider the performance of the widely adopted K-means clustering algorithm when the classification variables are correlated. We measure performance in terms of recovery of the true data structure. As expected, performance worsens considerably if the groups have elliptical instead of spherical shape. We suggest some modifications to the standard K-means algorithm which considerably improve cluster recovery. Our approach is based on a combination of careful seed selection techniques and use of Mahalanobis instead of Euclidean distances. We show that our method performs well in a number of examples where the standard algorithm fails. In such applications our nonparametric technique is seen to be competitive when compared to parametric model-based clustering methods. Hence, our conclusion is that use of the Mahalanobis distance should become a standard option of the available K-means routines for non-hierarchical cluster analysis. This goal can be achieved by minor modifications in popular commercial software.