Clustering algorithm based on Condensed Set Dissimilarity for high dimensional sparse data of categorical attributes

Sen Wu, Juanjuan Liu, Guiying Wei · 2011

Categorical data clustering is always challenging, especially when data is high dimensional and sparse. This paper proposes a new algorithm, named as CABOC, for clustering high dimensional sparse data with categorical attributes. Based on a new defined concept `Condensed Set Dissimilarity', the algorithm computes the dissimilarity of all the objects with sparse categorical attributes in a set directly. Furthermore, the algorithm only records a Condensed Set Reduction vector of the set during the computation process, which is defined to simply and accurately represent the necessary information of all the objects with sparse categorical attributes in the set for the clustering. So the computational complexity of the algorithm is low. A numeric example for customer cluster analysis illustrates the effectiveness of the algorithm.

Read the paper · More papers on PaperTik