Clustering Categorical Data

Yi Zhang, Ada Wai-Chee Fu, Chun Hing Cai, Pheng‐Ann Heng · 2005

Clustering has typically been a problem related to numerical data. However, in databases, oftentimes the data values are categorical and cannot be assigned meaningful numerical substitutes. With the recent interest in data mining, we begin to question the possibility of clustering numerical data. Following some recent work in this area, we propose an algorithm based on dynamical systems. To our knowledge, this is the first such algorithm that can guarantee the convergence of the dynamical system, which is a very important property for successful application. We demonstrated the effectiveness of the proposed method on both real data and synthetic data. We also propose a second method based on a graph partitioning approach, for which a new definition of similarity between two nodes is tailored for categorical data. 1 Introduction Mining numerical data has received much attention in recent research in data mining. One important form of knowledge that can be derived from such da...

Read the paper · More papers on PaperTik