A SURVEY: HIERARCHICAL CLUSTERING ALGORITHM IN DATA MINING

Seema Maitrey · 2012

The vast amount of information hidden in huge a database has created tremendous interests in the field of data mining. Among several data mining techniques, one of them called clustering is discussed in this paper. Clustering is the unsupervised classification of patterns (observations, data items, or feature vectors) into groups (clusters). A cluster is therefore a collection of objects which are similar between them and are dissimilar to the objects belonging to other clusters. Various clustering techniques available based on different parameters like distance, density, hierarchy and partition. The clustering problem has been addressed in many contexts and by researchers in many disciplines; this reflects its broad appeal and usefulness as one of the steps in exploratory data analysis. Whether for understanding or utility, cluster analysis has long been used in a wide variety of fields: psychology and other social sciences, biology, statistics, pattern recognition, information retrieval, machine learning, and data mining. The scope of this paper is modest: to provide an introduction to hierarchical clustering algorithm in the field of data mining, where we define data mining to be the discovery of useful, but non-obvious, information or patterns in large collections of data. A number of hierarchical clustering methods that have recently been developed are described here, with a goal of providing useful recommendation and references to fundamental concepts accessible to the broad community of clustering

Read the paper · More papers on PaperTik