Improved density peak clustering for large datasets

Vincent Courjault-Rade, Ludovic d’Estampes, Stéphane Puechmorel · HAL (Le Centre pour la Communication Scientifique Directe) · 2016

Clustering is the usual way of classifying data when there is no a priori knowledge,especially about the number of classes. Within the frame of big data analysis, thecomputational effort needed to perform the clustering task may become prohibitiveand motivated the construction of several algorithms or the adaptation of existing1ones, as the well known K-means algorithm . Recently, Rodriguez and Laio proposed an algorithm that clusters efficiently by fast searching local density peaksthat are sufficiently distant one from the others. However it is able to work on smalldatasets only and is highly sensitive to the value of tunable parameters. In this paperwe propose Improved Density Peak Clustering (IDPC), a new algorithm designed forlarge datasets based on [17] which corrects the shortcomings mentioned above. Thanksto our Cover Map (CM) procedure iterated with a decreasing locally-adaptive window(ICMDW), we are able to build both a localisation map and a multidimensionaldensity map. The nature of the density map, which fits perfectly with the approachof [17], allows us to compute the different steps with much less operations. It carriesunsensitive parameters, supports last improvements on cluster centers selection andpotentially allows new improvements.

Read the paper · More papers on PaperTik