Research and application of cluster analysis algorithm

Hailong Chen, Chunli Liu · 2013

With the popularity of database technology matures and data applications, the amount of data accumulated by the human increases rapidly. Facing the extremely large amount of data, we gradually step into a “rich data, poor knowledge” embarrassing situation. The data mining (Data Mining) rise to solve this problem. In this paper, we study the means and methods of clustering analysis that processing data partition or grouping, which is an important field in data mining. Based on the understanding of theoretical basis of clustering analysis, firstly, analyze in detail main algorithms of partitioning methods, hierarchical methods, density-based methods, grid-based methods and model-based methods. Secondly, compare performance of different clustering algorithms from scalability, the shape of cluster, sensitivity to the “noise”, and sensitivity to the data input sequence, high dimension and algorithm efficiency. Finally, use MATLAB for simulating and verifying applications of the algorithms based on K-means clustering analysis and hierarchical clustering.

Read the paper · More papers on PaperTik