KP-S: A Spark-Based Design of the K-Prototypes Clustering for Big Data

Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N’Cir, Nadia Essoussi · 2017

Big data is often characterized by a huge volume and a mixed types of attributes namely, numeric and categorical. K-prototypes is one of the most well-known clustering methods to deal with mixed data. Several parallel alternatives based on MapReduce have been proposed to enable this method to handle large scale of mixed data. However, these solutions are not suitable when dealing with Big data, due to time and memory restrictions. To address this issue, we propose in this paper a new Spark-based k-prototypes clustering method which uses the reclustering technique. We take advantage of the in-memory operations of Spark to build grouping from large scale of mixed data. Experiments performed on simulated and real data sets show that the proposed method is scalable and improves the efficiency of the existing k-prototypes methods.

Read the paper · More papers on PaperTik