Integrative biclustering of heterogeneous datasets using a Bayesian nonparametric model with application to chemogenomics
Dazhuo Li, Eric Christian Rouchka · BMC Bioinformatics · 2011
We present the Weighted Infinite Relational Model (WIRM) that jointly detects biologically sensible ligand groups and protein groups by integrating the clustering of various data types including chemical compound descriptors, protein sequences, ligand-target bindings and pharmaceutical effects. WIRM takes advantage of the Bayesian nonparametric paradigm for integrating multiple data types, for allowing for missing values (e.g. unknown ligand-target interaction) in the data, for automatically inferring the number of clusters without explicit model comparison, and for predicting the ligand-target interactions. Because some of these data types, to varying degrees, may suggest relationships having no implication for ligand-target interactions or for biological sensible ligand and protein groups, WIRM allows different types of data to have different weights based on prior knowledge of their quality or relevance. We apply WIRM to the ion channel proteins and G-protein-coupled receptors. We validate its performance using functional annotation and ligand-target interaction. We also test the relationship among multiple data types by varying the weights which indicate the impact of each data type on the model. The categories and interactions inferred by WIRM both confirm known biology and suggest novel predictions.