An improvement to the K-means algorithm oriented to big data

Joaquín Pérez-Ortega, R. Pazos, Miguel Hidalgo, Nelva Nely Almanza-Ortega, Ocotlán Díaz-Parra, René Santaolaya, Vitervo López Caballero · AIP conference proceedings · 2015

The K-means clustering algorithm is widely used in several domains, because of its simplicity of implementation and interpretation. However, one of its limitations is its high computational complexity. In this work the problem of reducing the complexity of the K means algorithm is approached, in order to make possible the solution of large scale data sets like those from Big Data, without significantly degrading solution quality. To this end, a new metaheuristics is proposed, which by an early assignment of objects to clusters, significantly reduces the number of calculations of distances from objects to centroids. The approach was experimentally evaluated by solving real and synthetic datasets yielding encouraging results. Time reductions of up to 91% were obtained with respect to the standard K-means, at the expense of reducing quality by 3.2%.

Read the paper · More papers on PaperTik