Application of new clustering algorithm designed for large data sets

Hua Dan-yang · Journal of Fuyang Teachers College · 2011

Clustering is one of important fields of research in data mining and is under vigorous development in recent years.With the amount of data doubling every three years,the key problem is how to do clustering on large data set efficiently and effectively.This paper focuses on application of current clustering analysis for large data sets,points out the key problem and technical challenges of existing clustering algorithms and presents proposal of solutions.Using a data structure DA-Tree(Data Aggregation Tree) as a common representation of large data sets,we have designed a new clustering algorithm(CLUK) and proved that it is superior to the existing algorithm(BIRCH).

Read the paper · More papers on PaperTik