Analysis of Clustering Algorithm Based on Number of Clusters, Error Rate, Computation Time and Map Topology on Large Data Set
Amanpreet Kaur Toor, Amarpreet Singh · 2013
Data clustering is an important task in the area of data mining. Clustering is the unsupervised classification of data items into homogeneous groups called clusters. Clustering methods partition a set of data items into clusters, such that items in the same cluster are more similar to each other than items in different clusters according to some defined criteria. Clustering algorithms are computationally intensive; particularly when they are used to analyze large amounts of data. With the development of information technology and computer science, high-capacity data appear in our lives. In order to help people analyzing and digging out useful information, the generation and application of data mining technology seem so significance. Clustering is the mostly used method of data mining. Clustering can be used for describing and analyzing of data. In this paper, the approach of Kohonen SOM and K-Means and HAC are discussed. After comparing these three methods effectively and reflect data characters and potential rules syllabify. This work will present new and improved results from large-scale datasets