Correction: A guide to bayesian networks software for structure and parameter learning, with a focus on causal discovery tools

Francesco Canonaco, Joverlyn D. Gaudillo, Nicole Astrologo, Fabio A. Stella, Enzo Acerbi · Frontiers in Systems Biology · 2025

Bayesian networks (BNs) have established themselves over the years as a powerful framework for modeling and analyzing complex systems under conditions of uncertainty. They have been widely employed in fields such as medicine (Arora et al., 2019), biology (Needham et al., 2007) and engineering (Kammouh et al., 2020). BNs represent probabilistic relationships among variables in a graphical way that allows efficient inference and intuitive causal reasoning when specific assumptions are met. It is important to clarify that while Bayesian networks encode conditional dependencies through directed edges, these do not necessarily imply causal relationships. A causal network is a specific type of Bayesian network where the edges reflect actual causal influences among variables, and their interpretation relies on assumptions such as causal sufficiency, faithfulness, and the absence of unmeasured confounding. Throughout this paper, we include structure learning algorithms developed for both probabilistic modeling and causal discovery. For a detailed discussion of the assumptions underlying causal discovery, we refer the reader to (Vonk et al., 2023). A BN Figure 1. Student Bayesian Network example with CPDs. (Jensen and Nielsen, 2007) • A collection of random variables represented as nodes X = {X 1 , X 2 , . . . , X n }, connected by directed edges that form a Directed Acyclic Graph (DAG). For instance, in Figure 1, the variables could be denoted as D (Difficulty), I (Intelligence), G (Grade), S (SAT), and L (Letter), corresponding to the nodes shown in the DAG.• A finite set of mutually exclusive states associated with each random variable.• For each random variable X i with parents Pa(X i ) = {Y 1 , . . . , Y n }, a Conditional Probability Distribution (CPD) specifying the probability distribution P (X i | Y 1 , . . . , Y n ). This CPD quantifies the influence of the parent variables on X i . If X i has no parents, it is associated with an unconditional probability distribution P (X i ). In Figure 1 Intelligence, while the recommendation Letter is assumed to be based exclusively on the Grade. This structure reflects the intuitive idea that each variable is directly influenced only by its parent nodes in the network (Koller and Friedman, 2009).In fact, BNs leverage conditional independence to compactly represent the joint probability distribution over a set of random variables X = {X 1 , X 2 , . . . , X n }. The joint distribution can be factorized into a product of CPDs, one for each node:P (X 1 , X 2 , . . . , X n ) = n i=1 P (X i | Pa(X i ))where Pa(X i ) denotes the set of parent variables of X i in the network. Although, in principle, various types of distributions can be used, most applications in the literature have focused on two main modeling assumptions due to their mathematical tractability and computational efficiency:• Discrete Bayesian Networks (Heckerman et al., 1995): assume that X i is a multinomial random variable dependent on the configurations of the values of its parents;• Gaussian Bayesian Networks (Geiger and Heckerman, 1994): assume that each variable X i is a univariate normal random variable, with its value linearly dependent on its parent variables.The objective of the learning process is to determine both the network structure and the associated parameters that best represent the observed data. Learning a Bayesian Network involves:• Structure learning: identifying the qualitative structure of the network, i.e., the conditional independence relationships among the variables.• Parameter learning: estimating the conditional probability distributions (CPDs) for each node.Learning the structure of a BN from data is a foundational step of the model construction process. For this purpose, a multitude of algorithms have been developed over the years; these methods are typically categorized into three groups: constraint-based, score-based, and hybrid.Constraint-based algorithms rely on the theory of causal graphical models introduced by Pearl (Verma and Pearl, 1990). A well-known example of this class is the PC-Stable (named after its authors Peter and Clark) algorithm (Colombo et al., 2014), which improves the original PC algorithm (Spirtes et al., 2000) by making it more robust to variable ordering. The algorithm starts with a complete undirected graph and recursively removes edges using a conditional independence (CI) test. Score-based algorithms define a scoring function, such as BIC (Bayesian information criterion) (Neath and Cavanaugh, 2012), AIC (Akaike information criterion) (Cavanaugh and Neath, 2019), to evaluate how well a given network fits the data.A search algorithm, such as greedy search or simulated annealing, is then used to explore the space of possible graphs. Hybrid algorithms combine constraint-based and score-based approaches. Typically, a constraint-based method is used to reduce the search space, followed by a score-based optimization over the reduced space.These algorithms generally assume that the input is tabular data, where each row represents an independent observation (i.i.d.), and each column corresponds to a variable. Constraint-based methods require data that are suitable for conditional independence (CI) testing, which typically includes discrete or continuous variables depending on the CI test used (e.g., chi-square for discrete, partial correlation for continuous). Score-based methods, on the other hand, rely on likelihood-based scoring functions and can handle discrete, continuous, or mixed data types depending on the scoring function and underlying assumptions. Hybrid methods inherit the data requirements of both approaches.To speed up or improve structure learning, prior knowledge can be incorporated to constrain or guide the search for the network structure. Users may specify relationships that are known to exist, permitted, or prohibited, thereby reducing the search space and enhancing both the accuracy and efficiency of learning algorithms.An overview of structure learning approaches is beyond the scope of this document; a comprehensive assessment of state-of-the-art methodologies can be found in (Nogueira et al., 2022;Kitson et al., 2023;Glymour et al., 2019;Scanagatta et al., 2019). Moreover, readers interested in the performance of the different classes of algorithms can refer to dedicated publications that offer comprehensive evaluations of the accuracy and computational efficiency of structure learning methods (Scutari et al., 2019).Parameter learning is another critical task in BNs development. Given the DAG, the objective of parameter learning is to estimate the parameters of the conditional probability distributions associated with each node, which is essential for inference and prediction. For a comprehensive review of parameter learning strategies, challenges, and algorithms, refer to the works of (Ji et al., 2015;Heckerman, 1998).Approaching the study of BN framework requires a solid understanding of fundamental principles in disciplines such as probability and computer science. Assuming that the reader is already familiar with these foundations, some convenient readings on causality and BNs science are offered by Probabilistic Graphical Models Principles and Techniques (Koller and Friedman, 2009), Bayesian Artificial Intelligence (Korb and Nicholson, 2010), Probabilistic Reasoning in Intelligent Systems (Pearl, 2014), Bayesian Networks with Examples in R (Scutari and Denis, 2021), Bayesian Networks in R with Application in the field of System Biology (Scutari and Lebre, 2013), Bayesian Networks and Influence Diagrams (Kjaerulff and Madsen, 2008). This document assumes that the reader is equipped with the necessary foundational knowledge and is ready to engage in practical hands-on work.Over the past five years, the field of causality and BNs development has seen an influx of numerous packages with no single solution being able to cater to all requirements and scenarios; this abundance of options is often challenging for individuals trying to gain hands-on experience with BNs. This document simplifies structure and parameter learning in BNs by providing a comprehensive overview of available software packages with a focus on causal discovery. In addition, we offer our subjective recommendations on selecting the best tools based on the reader's specific objectives. The remainder of this paper is structured as follows: Section 2 provides a systematic review of both open-source and commercial software. Section 3 offers guidance on selecting tools suitable for beginners. Section 4 summarizes the key contributions of this work. A concise summary of all reviewed tools is provided in Table S1 (Supplementary Material).2.1 gCastle gCastle (Zhang et al., 2021) is an end-to-end Python toolbox created by Huawei Noah's Ark Lab for causal structure learning. The package is equipped with functionalities such as data generation from simulated or real-world datasets, causal structure learning, and evaluation metrics.bnlearn (Scutari, 2010) Pgmpy (Ankan and Panda, 2015) is a Python library developed in 2015 by Ankur Ankan to work with probabilistic graphical models. It allows users to create their graphical models and then perform inferences or map queries to them. The library implements several inference algorithms like variable elimination, belief propagation, etc. The library is designed with a modular structure, allowing users to access dedicated classes for commonly used graphical models like Naive Bayes (NB) and hidden Markov models, eliminating the need to build them from base models. Currently, it includes implementations of various algorithms for structure learning, parameter estimation, both approximate, i.e., sampling-based, and exact inference, as well as causal inference.Tetrad (Ramsey et al., 2018) is a Java suite of software for the discovery, estimation, and simulation of causal models developed by the Carnegie Mellon University-Causal Learning and Reasoning (CMU-CLeaR) group. Some of its basic features for beginners include the ability to load existing datasets, load existing causal graphs, and create a new causal graph. For practitioners, the tool is equipped with advanced functionalities, such as specifying prior knowledge on constraint-based algorithms, manipulating data by imputing missing values, discretizing data, simulating data from statistical models, and computing the probability distribution of any variable, among others. It features a graphical user interface (GUI) and offers popular constraint-based algorithms for causal discovery such as PC, Fast Causal Inference (FCI), PC-Max, Conservative PC (CPC), and MLE for parameter learning.Causal-cmdfoot_1 is a Java application that offers a command-line interface tool for causal discovery algorithms developed by the Center for Causal Discovery. Currently, the application includes more than 30 algorithms for causal discovery. CDT (Kalainathan et al., 2020) is a Python package for causal inference in graphical models and pairwise settings (compatible with Python ≥ 3.5). Developed by Diviyan Kalainathan and Olivier Goudet, CDT provides tools for structure learning and dependency analysis. It leverages on NumPy, scikit-learn, PyTorch, and R to implement various algorithms for causal discovery, including methods from bnlearn and pcalg.The package is particularly suited for analyzing observational data, offering both classical and deep learning-based approaches to causal structure recovery.pyAgrum (Ducamp et al., 2020) is a Python wrapper for the C++ aGrUM library. It offers a high-level interface to aGrUM, enabling users to create, model, learn, apply, compute, and integrate BNs and other graphical models. Some specific (Python and C++) codes are added to simplify and extend the aGrUM API. The package contains causal discovery, parameter learning, and inference algorithms.Bnlearnfoot_2 is a Python package for causal discovery, parameter learning and inference developed by Erdogan Taskesen. It implements the most classical approaches for causal discovery such as HC, exhaustive search, Chow-Liu, TAN, PC, and MLE, as well as Bayesian estimation for parameter learning.OpenMarkov (Arias et al., 2019) is a Java open-source software tool developed by the Research Centre for Intelligent Decision-Support Systems. OpenMarkov comes with a user interface and can perform causal discovery employing the PC algorithm and HC search.Pomegranate (Schreiber, 2018), a Python package developed by Jacob Schreiber, offers efficient and versatile probabilistic models, spanning from individual probability distributions to composite models including BNs and hidden Markov models. The package offers both constraint-based and score-based algorithms, as well as parameter learning procedures.BayesFusionfoot_3 is a commercial software offering different solutions for causal discovery, parameter learning, and inference. Their flagship product is GeNIe, a tool for artificial intelligence and machine learning that has at its core the BN framework and other types of graphical probabilistic models. The SMILE engine allows the user to include custom applications that can be written in a variety of programming languages, e.g., C++, Python, Java, .NET, R, Matlab. Models created with GeNIe or SMILE can be shared or used on mobile devices via BayesMobile, or through a web browser with BayesBox.BayesiaLabfoot_4 is a commercial software developed by Dr. Lionel Jouffe and Dr. Paul Munteanu and their team. It offers plenty of algorithms for causal discovery, parameter learning, and inference. The software includes a graphical user interface and is well documented.Bayes Serverfoot_5 is a commercial software developed by Bayes Server Ltd. Besides the most well-known algorithms for causal discovery, parameter learning and inference, the software offers a wide range of tools for diagnostic, anomaly detection and decision-making under uncertainty which have at their core the BN framework. Bayes Server can be used in the cloud as well as on a local machine through a GUI. It offers an advanced user interface accessible programmatically via a number of APIs that can be used via Java, Matlab, Python, Spark and R.This section aims to assist beginners select the ideal package or software that best suits their needs. The first subsection focuses on causal discovery tools, while the second presents tools that support functionalities for both parameter learning and structure learning for the Bayesian network framework. Finally, the last subsection discusses commercial software that offers additional features such as optimized user-interfaces and professional customer support. Note that while the previous section provided a comprehensive overview of available solutions, this section shortlists and discusses only those we consider most suitable for beginners.It is important to note that while all the tools discussed in this section aim to uncover structure among variables, they differ in their underlying modeling assumptions and output types. Some tools (e.g., bnlearn, pgmpy, pyAgrum) are focused on Bayesian networks and provide probabilistic modeling capabilities, including structure and parameter learning as well as inference. Others (e.g., LiNGAM, CDT, causal-learn) are specialized for causal discovery and do not build a full probabilistic graphical model. Instead, these methods aim to recover a causal DAG under specific assumptions (e.g., linearity, non-Gaussianity, no hidden confounding). While the outputs may look similar (DAGs), their interpretation and use cases are different. We highlight these distinctions the section to readers select the tool that best fits their the is to the underlying structure among variables typically under assumptions the need for full probabilistic modeling or inference, CDT, and are three tools that represent solutions and provide access to those gCastle by Huawei Noah's Ark Lab is in our one of the most accessible and comprehensive causal discovery open-source Python at the of this It offers various approaches for the structure of causal networks from score-based to and For each algorithm, the offers a detailed practical making the tool to beginners. can be found in Causal Inference and in Python Causal which offers the user the ability to into any offered by the Moreover, gCastle can be used via a which provides a of the interface that not CDT is another package that we in contains several that guide users into their first learning CDT has the collection of algorithms for causal discovery among all the other reviewed tools for some of which can be using as data, the model et al., the original framework et al., to for It assumes that each variable is a function of its past values and the past values of other variables, a number of The model assumes that the are continuous, and independent over and independence assumptions are essential for identifying the of causal relationships from observational data, which be under Gaussian Python package includes implementations for various models, including the model for It offers and practical for each model, making it a tool for both and causal the is causal discovery on data, represents a This Java library implements several algorithms for causal discovery and can be used via a or as of a We this library to be to the we a for more or advanced our assessment of tools specialized in causal learning, we consider CDT to be the best when a set of available methodologies is For CDT could be the most for or where and the of various methods is CDT is the best when an interface with is or While CDT offers a wide range of causal discovery algorithms, gCastle for its interface and making it accessible to both CDT and gCastle support models, the most suitable tool when with this model as it is for such one may to both the structure and the parameters of a probabilistic model using the Bayesian network framework. this several tools both while offering of of the most complete and tools to is from the of methods for parameter learning, learning, inference, missing data and model strategies, bnlearn is its and practical most methods and are in the Bayesian Networks in R and Bayesian Networks Examples in R (Scutari and Denis, 2021), of which the of bnlearn is to bnlearn is represented by bnlearn, which provides methods for the the as This is an important given that a of real-world and systems include the other hand, the range of algorithms available in is more than in bnlearn, particularly for structure learning for this by offering more comprehensive are available in the practical with both of which are for the first into this to bnlearn is like pgmpy, provides methods for and making it a for real-world offers comprehensive including and applications with important offered by is a of solutions to the in the of by not only provides and offers a wide of structure learning For it implements greedy local search with Chow-Liu, TAN, and package is the Python of the original bnlearn is an R it is not as in methodologies as Pgmpy and the original bnlearn a of causal discovery algorithms are available in the Python of bnlearn offers an intuitive interface and its is as and as the original R The not only presents followed by the associated output provides a to the theory for those are familiar with R, bnlearn represents the best when with the For Python need to model and are the best to the multitude of in and provides added value for beginners their first in this The Python of bnlearn offers a more interface than the other it with a number of structure learning algorithms, making it suitable for readers to with a and more application of BNs in settings where cloud computing be the models often need to be shared and from a variety of including mobile where solutions may be In addition, in these of professional support is making the open-source packages in the previous In this we some practical commercial solutions that the of this purpose, Bayes Server be our A is available on their The provides comprehensive including how to with the graphical user The features a section that as a of practical for with the Bayes Server API. the numerous real-world use cases various including and Server is available under both commercial and by is a to Bayes GeNIe use of the SMILE a library of C++ classes that implement causal and parameter learning, as well as inference, which can be via API. SMILE can be used via Java, Python, R, and using the and that can and use the (Python and of GeNIe is an where graphical models can be and from a variety of including of is available on the provides detailed which includes information GeNIe and its main as well as and for The support is and can be a for to and GeNIe is has a commercial and offers an intuitive and such as an that includes several and use cases that users the multitude of features offered by are It is to that the framework can be using Java both and GeNIe can the They are both equipped with a web that features a interface and and both software can be used on mobile For and GeNIe, and models can be the in which tool best suits the reader's after their and This not to as users the software on the can only be used with paper provides an overview of tools and software packages for Bayesian network structure and parameter learning, as well as methods developed for causal discovery. The tools reviewed from the of a to gain hands-on experience in the and subjective recommendations given which tools are more the it is important to that the of BN tools This is due to the range of data types (e.g., discrete, continuous, and application (e.g., that BN modeling a packages have been developed to cater to specific to and a of we the field is a methodologies that beyond are in real-world be to integrate the software in this paper into more and as like et al., various machine learning algorithms into a we the of for BN modeling that with not only and the of for and Given the of this of this document be

Read the paper · More papers on PaperTik