A Novel K-Means Based Clustering Algorithm for High Dimensional Data Sets

Madjid Khalilian, Norwati Mustapha, Nasir Suliman, Ali Mamat · 2010

Abstract- Data clustering is an unsupervised method for extraction hidden pattern from huge data sets. Having both accuracy and efficiency for high dimensional data sets with enormous number of samples is a challenging arena. For really large and high dimensional data sets, vertical data reduction should be performed prior to applying the clustering techniques which is performing dimension reduction, and the main disadvantage is sacrificing the quality of results. However, because dimensionality reduction methods inevitably cause some loss of information or may damage the interpretability of the results, even distorting the real clusters, extra caution is advised. Furthermore, in some applications data reduction is not possible e.g. personal similarity, customer preferences profiles clustering in customer recommender system or data by which is generated sensor networks. Existing clustering techniques would normally apply in a large space with high dimension; dividing big space into subspaces horizontally can lead us to high efficiency and accuracy. In this study we propose a method that uses divide and conquer technique with equivalency and compatible relation concepts to improve the performance of the K-Means clustering method for using in high dimensional datasets. Experiment results demonstrate appropriate accuracy and speed up.

Read the paper · More papers on PaperTik