Parallel clustering method for non-disjoint partitioning of large-scale data based on spark framework
Zayani Abir, Chiheb-Eddine Ben N’Cir, Nadia Essoussi · 2016
Clustering large scale data has become an important challenge which motivates several recent works. While the emphasis has been on the organization of massive data into disjoint groups, this work considers the identification of non-disjoint groups rather than the disjoint ones. In this setting, it is possible for data object to belong simultaneously to several groups since many real-world applications of clustering require non-disjoint partitioning to fit data structures. For this purpose, we propose the Parallel Overlapping k-means method (POKM) which is able to perform parallel clustering processes leading to non-disjoint partitioning of data. The proposed method is implemented within Spark framework to ensure the distribution of works over the different computation nodes. Experiments which we have performed on simulated and real-world multi-labeled datasets shows both faster execution times and high quality of clustering compared to existing methods.