Clustering Categorical Data Based on Representatives
S. Aranganayagi, K. Thangavel · 2008
Clustering of categorical data is one of the data mining techniques, which helps in identifying clusters within the domain space. In this paper we present a new method to cluster categorical data. This new representative based method works in three phases. The dissimilarity matrix, neighbor matrix and the initial clusters are formed in first phase. Merging of clusters is performed in the second phase by relocating the objects using the neighborhood concept. In the third phase, mode of attributes of clusters is computed, and phase I and Phase II are applied for the tuples formed from these representatives. The proposed method is experimented with the well known data sets from UCI data repository, soybean, zoo and mushroom data set.