MapReduce-based k-prototypes clustering method for big data

Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N’Cir, Nadia Essoussi · 2015

Big data clustering is one of the recently challenging tasks that is used in many application domains. Traditional clustering methods are not able to deal with large-scale of data. Furthermore, Big data are often characterized by the mixed type of data, including numerical and categorical attributes. Thus, we propose in this paper the parallelization of k-prototypes clustering method (MR-KP) using MapReduce model to handle large-scale of mixed data. Experiments results show that MR-KP scales well with increasing data set sizes and achieves a close to linear speedup while maintaining the clustering accuracy.

Read the paper · More papers on PaperTik