Hybrid hierarchical clustering: Forming a tree from multiple views
Anjum Gupta, Sanjoy Dasgupta · 2005
We propose an algorithm for forming a hier-archical clustering when multiple views of the data are available. Different views of the data may have different underlying distance mea-sures which suggest different clusterings. In such cases, combining the views to get a good clustering of the data becomes a challenging task. We allow these different underlying dis-tance measures to be arbitrary Bregman di-vergences (which includes squared-Euclidean and KL distance). We start by extending the average-linkage method of agglomerative hierarchical clustering (Ward’s method) to accommodate arbitrary Bregman distances. We then propose a method to combine mul-tiple views, represented by different distance measures, into a single hierarchical cluster-ing. For each binary split in this tree, we consider the various views (each of which suggests a clustering), and choose the one which gives the most significant reduction in cost. This method of interleaving the dif-ferent views seems to work better than sim-ply taking a linear combination of the dis-tance measures, or concatenating the feature vectors of different views. We present some encouraging empirical results by generating such a hybrid tree for English phonemes. 1.