A Contrast Pattern Based Clustering Algorithm for Categorical Data
Neil Fore · OhioLink ETD Center (Ohio Library and Information Network) · 2010
A Contrast Pattern based Clustering Algorithm for Categorical Data.The data clustering problem has received much attention in the data mining, machine learning, and pattern recognition communities over a long period of time.Many previous approaches to solving this problem require the use of a distance function.However, since clustering is highly explorative and is usually performed on data which are rather new, it is debatable whether users can provide good distance functions for the data.This thesis proposes a Contrast Pattern based Clustering (CPC) algorithm to construct clusters without a distance function, by focusing on the quality and diversity/richness of contrast patterns that contrast the clusters in a clustering.Specifically, CPC attempts to maximize the Contrast Pattern based Clustering Quality (CPCQ) index, which can recognize that expertdetermined classes are the best clusters for many datasets in the UCI Repository.Experiments using UCI datasets show that CPCQ scores are higher for clusterings produced by CPC than those by other, well-known clustering algorithms.Furthermore, CPC is able to recover expert clusterings from these datasets with higher accuracy than those algorithms.