Bayesian Partitioning of Large-Scale Distance Data

David Adametz, Volker Röth · 2011

A Bayesian approach to partitioning distance matrices is presented. It is inspired by the Translation-invariant Wishart-Dirichlet process (TIWD) in [1] and shares a number of advantageous properties like the fully probabilistic nature of the in-ference model, automatic selection of the number of clusters and applicability in semi-supervised settings. In addition, our method (which we call fastTIWD) over-comes the main shortcoming of the original TIWD, namely its high computational costs. The fastTIWD reduces the workload in each iteration of a Gibbs sampler from O(n3) in the TIWD to O(n2). Our experiments show that the cost reduction does not compromise the quality of the inferred partitions. With this new method it is now possible to ‘mine ’ large relational datasets with a probabilistic model, thereby automatically detecting new and potentially interesting clusters. 1

Read the paper · More papers on PaperTik