A CLUSTERING ALGORITHM FOR MIXED NUMERIC AND CATEGORICAL DATA

OhnMarSan, Van‐Nam Huynh, YoshitexuNakamori · 系统科学与复杂性:英文版 · 2003

Most of the earlier work on clustering mainly focused on numeric data whose inherent geometric properties can be exploited to naturally define distance functions between data points. However, data mining applications frequently involve many datasets that also consists of mixed numeric and categorical attributes. In this paper we present a clustering algorithm which is based on the k-means algorithm. The algorithm clusters objects with numeric and categorical attributes in a way similar to k-means. The object similarity measure is derived from both numeric and categorical attributes. When applied to numeric data, the algorithm is identical to the k-means. The main result of this paper is to provide a method to update the 'cluster centers' of clustering objects described by mixed numeric and categorical attributes in the clustering process to minimise the clustering cost function. The clustering performance of the algorithm is demonstrated with the two well known data sets, namely credit approval and abalone databases.

Read the paper · More papers on PaperTik