Clustering based on Dissimilarity First Derivatives
Ana L. N. Fred · Pattern Recognition in Information Systems · 2002
A hierarchical agglomerative clustering algorithm based on the analysis of dissimilarity increments between neighboring patterns is presented. The first derivative of dissimilarity between neighboring patterns inside a natural cluster is modelled by an exponential distribution, this statistic characterizing the cluster. A cluster isolation criterion is defined based on estimates of each cluster dissimilarity increments mean value, continuously updated along the clusters formation process, under a hierarchical agglomerative framework. Unreliable estimates, mainly occurring when cluster cardinality is low, can lead to over-fragmentation of the data into spurious, small sized clusters. In order to prevent this situation, a regularizing function is proposed to widen the estimates of the exponential distribution mean, when the number of samples is small. Analysis of the method is performed in a comparative study with the well known single-link and k-means algorithms. Application examples using both syntectic and real data show the ability of the method to identify arbitrary shaped clusters.