Realization of K-means Clustering Algorithm Based on Hadoop
Luo Cheng · Computer Technology and Development · 2013
For the problem of high time complexity of K-means algorithm,propose a method using MapReduce programming model and Hadoop cloud platform to reduce the time complexity of K-means algorithm in dealing with huge data.Design Map function to calculate the distance of each record to each center key and mark their categories,and design Reduce function to update the center keys and calculate the distance of each record to its center key,then make a summary of the distance results.Through the experiment,verify that compared with the traditional serial algorithm when dealing with huge data,the new K-means algorithm can indeed reduce the time complexity,and also has good stability and expansibility.