Clustering and Validation for Very Large Databases (VLDB)

B. F. Momin · 2006

The digital revolution has made digitized information easy to capture and fairly inexpensive to store in database at exponential rate. Unfortunately, these very large databases (VLDB) are not able to analyze in a reasonable amount of time and cost. Clustering is one such operation to group similar objects based on their distance, connectivity, relative density or some specific characteristics. Predicting the correct number of clusters and its quality evaluation are the key issues. This paper review various techniques for clustering and its validation. It presents a framework for cluster validation with most commonly used validity indices. The experiments on microarray gene expression dataset demonstrate the clustering and its validations. The results obtained indicate how to choose quality cluster that helps in data mining.

Read the paper · More papers on PaperTik