The Projected Dip-means Clustering Algorithm

Theofilos Chamalis, Aristidis C. Likas · 2018

One of the major research issues in data clustering concerns the estimation of number of clusters. In previous work, the dip-means clustering algorithm has been proposed as a successful attempt to tackle this problem. Dip-means is an incremental clustering algorithm that uses a statistical criterion called dip-dist to decide whether a set of data objects constitutes a homogeneous cluster or it contains subclusters. The novel aspect of dip-dist is the introduction of the notion of data unimodality when deciding for cluster compactness and homogeneity. More specifically, the use of dip-dist criterion for a set of data objects requires the application of a univariate statistic hypothesis test for unimodality (the so called dip-test) on each row of the distance (or similarity) matrix containing the pairwise distances between the data objects. In this work, we propose an alternative criterion for deciding on the homogeneity of a set of data vectors that is called projected dip. Instead of testing the unimodality of the the row vectors of the distance matrix, the proposed criterion is based on the application of unimodality tests on appropriate 1-d projections of the data vectors. Therefore it operates directly on the data vectors and not on the distance matrix. We also present the projected dip-means (pdip-means) algorithm that is an adaptation of dip-means using the proposed pdip criterion to decide on cluster splitting. We conducted experiments using the pdip-means and the dip-means algorithms on artificial and real datasets to compare their clustering performance and provide empirical conclusions from the obtained experimental results.

Read the paper · More papers on PaperTik