Bayesian Networks for Gene Networks Discovery: Parallel and Optimised Learning
Marco Scutari · 2013
In genetics and systems biology, Bayesian networks (BNs) are used to describe and iden-tify interdependencies among genes and gene products, with the eventual aim to better understand the molecular mechanisms that link them. If we assign each gene to one node in the BN, edges represent the interplay between different genes, and can describe either direct (causal) interactions or indirect influences that are mediated by unobserved genes. BNs can be estimated (learned) with a variety of algorithms, which can all be traced to three approaches: 1. constraint-based, which are based on conditional independence tests; 2. score-based, which are based on goodness-of-fit scores; 3. and hybrid, which combine the previous two approaches. Score-based algorithms are just the application of general purpose optimisation techniques to BNs, and most are inherently sequential (e.g. each step depends on the previous one). On the other hand, constraint-based algorithms can be parallelised effectively, to the point that it is feasible to learn gene networks from high-dimensional data. Parallel Constraint-Based Learning