Evaluation of Clustering Algorithms for High Dimensional Data Based on Distance Functions

Smita Chormunge, Sudarson Jena · 2014

Clustering high-dimensional data spaces are often encountered in areas such as medicine, DNA analysis in computational biology and many others. It imposes on a data analysis severe computational requirements and present real challenges to clustering algorithms. This paper evaluates the performance efficiency of K-means and Agglomerative hierarchical clustering methods based on Euclidean and Manhattan distance functions for high dimensional data. Efficiency concerns the computational time required to build up datasets. Extensive experiments carried out to evaluate two clustering methods on Microarray datasets. The results demonstrate that Agglomerative hierarchical clustering algorithm is efficient in time for both distance functions than K-means clustering algorithm.

Read the paper · More papers on PaperTik