Faster Mahalanobis K-means clustering for Gaussian distributions

Ankita Chokniwal, Manoj K. Singh · 2016

The famous probabilistic theory of Gaussian distribution suggests that distribution of real world data when collected in large quantities always follows Gaussian distribution. The increasing amount of data to be managed and analyzed currently requires clustering approaches which are successful in recognizing the dense areas of Gaussian model. These areas might not be always spherical, as often discovered by k-means; rather the shapes might be oblong. The shapes of clusters formed highly depend on the distance metric used. Use of Mahalanobis distance to identify clusters in a mixed Gaussian distribution field has always been appreciated, reasons being Gaussian Mixture Models (GMMs) supporting Mahalanobis distance and its ability to identify real life-like elliptical clusters. The only challenge is deciding the proper initial estimates used in computation of Mahalanobis distance. The elliptical clusters recognized by use of this measure in k-means can be viewed as a general case of spherical clusters produced by Euclidean distance in k-means. This paper explores how well the initialization method of k-means++ works with Mahalanobis k-means.

Read the paper · More papers on PaperTik