Associative Clustering by Maximizing a Bayes Factor

Janne Sinkkonen, Janne Nikkilä, Leo Lahti, Samuel Kaski · 2003

Clustering by maximizing the dependency between (margin) groupings or partitionings of co-occurring data pairs is studied. We suggest a probabilistic criterion that generalizes discriminative clustering (DC), an extension of the information bottleneck (IB) principle to labeled continuous data. The criterion is the Bayes factor between models assuming dependence and independence of the two cluster sets, and it can be used as a well-founded criterion for IB for small data sets. With suitable prior assumptions the Bayes factor is equivalent to the hypergeometric probability of a contingency table with the optimized clusters at the margins, and for large data it becomes the standard mutual information. An algorithm for two-margin clustering of paired continuous data, associative clustering (AC), is introduced. Genes are clustered to find dependencies between gene expression and transcription factor binding, and dependencies between expression in di#erent organisms.

Read the paper · More papers on PaperTik