Comparative study of data mining clustering algorithms

Iyer Aurobind Venkatkumar, Sanatkumar Jayantibhai Kondhol Shardaben · 2016

In today's world, where we generate large amount of data, we can harness the benefits of the hidden information i.e. patterns or correlations in these data. This information can be used in various constructive fields only if we are able to handle big data efficiently. One such process that is used to extract and handle the hidden information is data mining. There are various techniques in data mining namely Clustering, Prediction, Classification, Association etc. Clustering is dividing data set into related groups such that all the groups do not have anything in common. Prediction, as the name suggests, predictions are made with available data set. It does not give surety of any kind, it may predict right or may predict wrong. Classification is classification of data sets into some predefined sets using various mathematical models. Association is discovering a correlation hidden in large amount of data, that is, in a given transaction based on the relationships between the items a pattern is discovered. In this paper we study one of the most widely used methods to handle big data, that is, data mining clustering algorithms. Here we have studied and made a comparative analysis of four classic clustering algorithms namely K-means, BIRCH, DBSCAN, STING.

Read the paper · More papers on PaperTik