Visual heuristics for data clustering

Tung-Duong Tran-Luu, Nicholas DeClaris · 2002

We are concerned with finding clusters in data by reordering the proximity matrix as close as possible into block diagonal form. We also define a new proximity measure for variables with word values that can be semantically consistent with our knowledge in the field in question. Moreover, we unify various measures of blockness into the form of a quadratic programming problem. We propose two new algorithms to reorder proximity matrices: MST linearization (MLin) and dendrogram linearization (DLin). Their performance is compared against four other popular algorithms by running numerous data sets, real and artificial. We find that MLin is competitive and complementary with the furthest neighbor, which is among the best existing algorithms.

Read the paper · More papers on PaperTik