Discussion of "learning equivalence classes of acyclic models with latent and selection variables from multiple datasets with overlapping variables"
Jiji Zhang, Ricardo Bezerra de Andrade e Silva · Digital Commons - Lingnan (Lingnan University) · 2011
In automated causal discovery, the constraint-based approach seeks to learn an (equivalence) class of causal structures (with possibly latent variables and/or selection variables) that are compatible (according to some assumptions, usually the causal Markov and faithfulness assumptions) with the conditional dependence and independence relations found in data. In the paper under discussion, Tillman and Spirtes (T&S) develop a constraint-based algorithm for learning causal structures from multiple, overlapping datasets. The basic setup of the problem is this: the variables of interest are not all measured at once in a single study. Instead there are several studies, each measuring a subset, which produce multiple datasets with overlapping variables. Assuming there is a common structure over the variables of interest (with possibly latent confounding variables and selection variables) that generated all the datasets, T&S’s algorithm is designed to discover features of that structure by learning the features shared by all the causal structures that are compatible with all the datasets. Unlike standard constraint-based methods, which assume an oracle of conditional independence that can respond to every query of whether two observed variables are conditionally independent given a set of observed variables, the algorithm described in T&S’s paper allows the oracle to be incomplete: some variables may be jointly measured in none of the available datasets; as a result, the query of conditional independence concerning these variables cannot be answered. Thus the algorithm provides a solution to the problem of learning from (a broad class of) incomplete oracles. To be sure, the algorithm does not allow the oracle to be arbitrarily incomplete: with respect to the variable set of each individual dataset, the oracle is still complete. But it seems to be general enough to handle the problem of incomplete oracle in some other contexts. For example, one may worry about the power of conditional inde-