A Simple Methodology for Database Clustering

Hao Tang, Zhang Mei · 2015

Database clustering is a preprocess technology for multi-database mining.Existing algorithms for database clustering are successful in terms of having a cluster, but the time complexity is high or excessive pursuit in respect of the non-trivial complete clustering, which may lead to a bad clustering or application-dependent.The practical application of these algorithms have high time instability and complexity.In this article, we put forward the application-independent database clustering methodology by using hierarchical clustering method to avoid instability and reduce time complexity.This methodology is called Database Hierarchical Clustering.We firstly construct a multi-objective optimization problem, and then use hierarchical clustering algorithm to find the optimal cluster structure thought.We also use the cophenetic correlation coefficient to evaluate the best cluster.Experiments on the synthetic databases and the real-world databases show that our method of clustering stability features lower time complexity than that of the BestClassification while also highlighting strong generalization ability.

Read the paper · More papers on PaperTik