Constructing non-parametric Bayesian networks from data

Stefan Blauw · Research Repository (Delft University of Technology) · 2014

The objective of this thesis is to design an algorithm for learning the structure of non-parametric Bayesian networks (BNs) from data and has been written as part of acquiring the bachelor's degree at the Delft Institute of Applied Mathematics of the TU Delft. Inspiration for improving an existing algorithm while designing a new procedure was provided by the study of Sneller in [11]. Non-parametric BNs can be used to model both continuous and hybrid data, but perform best when only continuous data is used. They do not require the severe assumption of the joint normal distribution, and there is no need to discretize the data. Instead, only the normal copula is assumed and has to be validated. This is done by comparing the determinant of the rank correlation matrix of computer generated data under the assumption of the normal copula (DNR) with the determinant of the rank correlation matrix of the actual data (DER). The BN starts with nodes only, and edges are added based on a high rank correlation (Spearman correlation). Colliders and paths between independent nodes are used to determine the direction of arcs. As this can only be done up to Markov equivalence classes, the last arcs are directed randomly, making sure that no cycles are created. Based on two public datasets a comparison is made with the well-known PC algorithm. It turns out that the number of edges added by the new algorithm is approximately 1.5 times higher than the number of edges added by the PC algorithm, while some nodes are left unconnected by the new algorithm when they were found connected by the PC algorithm. The procedure for directing edges on paths between independent nodes and for finding colliders has improved compared with Sneller. Using an artificial dataset that containes independent nodes, the algorithm modeled all but one independence relations well. It was not possible to entirely model the independencies correct, as another independence would have been violated. A few arcs that could not be learned from the data were directed randomly.

Read the paper · More papers on PaperTik