The Application of User Behavior Analysis by Improved K-means Algorithm Based on Hadoop
Revista de la Facultad de Ingeniería · 2016
In recent years, the new social network and mobile Internet technology promote the rapid growth of the number of network users, as well as the network data showes explosive growth.How to analysisand extract the users behavior characteristics from the mass data mining is more and more important. There are many clustering algorithms for user behavior analysis, and K-means algorithm is one of the most common method. However, there are some disadantages of K-means that lead to reduce the performance of clustering, which include that the number of clusters and the cluster center must be initialized, it issensitive to abnormal data, andit can only handle the numerical data.Based on the research of Hadoop, and clustering algorithms, we propose a K-means clustering method based on Canopy method.Finally, we carried out a comparative analysis of results of cluster experiment by comparing to the single machine experiment.The experimental results show that the K-means clustering algorithm based on Canopy is faster than the original K-means algorithm, which means that the algorithm has better expansibility, and it has a high application value.