A Review of Clustering Algorithms

Suchita S. Mesakar, Manoj S. Chaudhari · 2013

Data mining is the process of extracting meaningful data or knowledge from large amount of data. Clustering is the dynamic field of research in data mining. Data clustering is used in variety of applications like pattern matching, machine learning, image segmentation and information retrieval. The aim of clustering is to group data into clusters or groups, so that data in the same cluster are more similar to each other than to those in other clusters. There is large amount of data available in the database; fast retrieval of data from database is always required. So clustering the data will ease the task of retrieval of data from database. This paper presents an overview of various clustering algorithms used for clustering numerical and categorical data. Clustering is a data mining technique and it plays a vital role in classification of data. The database consists of large amount of data. This large amount of data can be grouped into meaningful data for further analysis or management. Clustering plays important role in management and analysis of data. Fast retrieval of data from database is always a need and if large amount of data is classified into meaningful groups or clusters then it will be easier and faster to access the data from the database. The goal of clustering is to cluster the data into groups or cluster depending on the similarity and dissimilarity measures. Data clustering is the process of organizing objects into groups whose members are similar in some way. Clustering algorithm partitions the data into certain number of clusters. Clustering is used in many areas like machine learning, pattern recognition together with data mining, document retrieval, image segmentation. Learning can be classified as supervised learning and unsupervised learning. Clustering can be considered as unsupervised learning problem as it deals with finding a structure in a collection of unlabeled data. In unsupervised learning for given set of patterns, a collection of clusters is to be discovered and additional patterns are assigned to correct cluster. In supervised learning set of classes (clusters) are given, new pattern (point) are assigned to proper cluster, and are labeled with label of its cluster. Clustering is often called an unsupervised learning task because no class values are given which denotes an a priori grouping of the data instances. Clustering is the dynamic field of research in data mining. There exist a large number of clustering algorithms in the literature. The clustering algorithms partition data into certain number of clusters based on similarity and dissimilarity. The choice of clustering algorithm depends both on the type of data available and on the particular purpose and application. The clustering can be performed on numerical data and categorical data. The numeric data can be ordered naturally and the properties can be used to apply distance measures to the attribute values. The examples of numeric attributes are age,cost,weight But in case of

Read the paper · More papers on PaperTik