Clustering Non-Ordered Discrete Data *
Alok Watve, Sakti K. Pramanik, Sungwon Jung, Bumjoon Jo, Sunil Kumar, Shamik Sural · 2014
Clustering in continuous vector data spaces is a well-studied problem. In recent years there has been a significant amount of research work in clustering categorical data. Howev-er, most of these works deal with market-basket type transaction data and are not specifi-cally optimized for high-dimensional vectors. Our focus in this paper is to efficiently cluster high-dimensional vectors in non-ordered discrete data spaces (NDDS). We have defined several necessary geometrical concepts in NDDS which form the basis of our clustering al-gorithm. Several new heuristics have been employed exploiting the characteristics of vec-tors in NDDS. Experimental results on large synthetic datasets demonstrate that the pro-posed approach is effective, in terms of cluster quality, robustness and running time. We have also applied our clustering algorithm to real datasets with promising results.