Survey of Recent Clustering Techniques in Data Mining
Anoop Jain, Satyam Maheswari · 2012
Cluster analysis or clustering is the task of assigning a set of objects into groups called clusters. Main task of clustering are explorative data mining, and a common technique for statistical data analysis used in many fields, including machine learning, pattern recognition, image analysis, information retrieval, and bioinformatics. Cluster analysis itself is not one specific algorithm, but the general task to be solved. It can be achieved by various algorithms that differ significantly in their notion of what constitutes a cluster and how to efficiently find them. Popular notions of clusters include groups with low distances among the cluster members, dense areas of the data space, intervals or particular statistical distributions. The appropriate clustering algorithm and parameter settings including values such as the distance function to use, a density threshold or the number of expected clusters depend on the individual data set and intended use of the results. Cluster analysis as such is not an automatic task, but an iterative process of knowledge discovery or interactive multi-objective optimization. It will often be necessary to modify preprocessing and parameters until the result achieves the desired properties. In this paper we represent a survey of clustering techniques in data mining. The clustering techniques are categorized based upon different approaches. This paper provides the major advancement in the clustering approach for data mining research using these approaches the features and categories in the surveyed work.