Parallel clustering algorithm based on sparse index sort of high dimensional data

Wu Qinghai · Systems Engineering - Theory & Practice · 2011

High dimensional data clustering is an important research subject of data mining.Large scale high dimensional data clustering is more challenging.A parallel algorithm P-CABOSFV is presented based on sparse index sort by extending CABOSFV,an efficient high dimensional data clustering algorithm and by using parallel computing pattern to improve the capability to deal with large scale data.The proposed algorithm partitions data according to high dimensional data sparse index sort,distributes the segmented data to some certain compute nodes to implement clustering tasks concurrently and then merges the results from all the nodes together to get the final clustering result depending on Sparse Feature Dissimilarities of the merged sets.Experiments using UCI and computer synthetic data sets show that P_CABOSFV algorithm has reliable clustering quality,strong scalability from data size and data dimensions,which means it is feasible and effective.

Read the paper · More papers on PaperTik