K-means+: A developed clustering algorithm for big data

Kun Niu, Zhipeng Gao, Haizhen Jiao, Nanjie Deng · 2016

Clustering is one of the most important task in data mining. But for big data application, clustering models are faced with the problem of high complexity for low respond time requirement. This paper focuses on velocity criterion of big data modeling, presents a developed k-means algorithm, k-means+, which effectively reduces time costs of clustering modeling through block operation and redesigning of distance function. Block operation aggregates instances as blocks to cluster afterwards. Manhattan distance is used instead of common Euclidean distance to simplify calculation. Experimental results show that k-means+ works well on most testing datasets and executes much faster than original k-means.

Read the paper · More papers on PaperTik