Research on Improved K-Means Algorithm Based on Hadoop
Wei Xiaojing, Yuanbo Li · 2017
The approaching of big data era makes traditional data storage methods could not accomplish the mass tasks of analysis, managing and mining. While the rise of cloud computing brings new life to parallel improvement of data mining algorithms. Its efficient programming model, massive storage capacity and powerful computing capabilities provide a broad platform for the development of data mining. Hadoop uses MapReduce programming model for distributed computing and HDFS distributed file system for file storage. This paper studied the K-means algorithm in the clustering algorithm in detail. Through the sampling of large-scale data, this paper used convex hull and the solution of the heel point to solve the initial two clustering centers to modify algorithm process. Through the MapReduce programming model, it achieved the entire process of parallelization. Finally, it compared the efficiency of the improved algorithm with different distance measurement methods, along with serial parallel mode and different cluster nodes in experiments. With the increasing of cluster nodes and data scale, reliability and computational efficiency of improved parallel algorithm are improved obviously.