A Study on Clustering Algorithms Used In Large Datasets
A.J. Anju, P. O. Sinciya · International journal of engineering and future technology · 2016
A collection of related sets of information that is composed of separate elements but can be manipulated as a unit by a computer. Data mining is the analysis of (often large) observational data sets to find unsuspected relationships and to summarize the data in novel ways that are both understandable and useful to the data owner .A Markov clustering algorithm that operates on a graph built from pair wise similarity information of the input data. Edge weights stored in the stochastic similarity matrix are alternately fed to the two main operations, inflation and expansion, and are normalized in each main loop to maintain the probabilistic constraint. The clustering is very difficult in terms of processing speed, time delay, search from different database etc. In hierarchical clustering the large scale data sets are difficult to capture, store, search, analyze and visualize. Hierarchical clustering cannot represent distinct clusters with similar expression patterns.This paper surveys different clusteringalgorithms which obtainedmaximum similarity of large protein sequence dataset.