When will causal structure learning become practical?
Johannes C. Textor · International Journal of Epidemiology · 2025
Epidemiologists use structural causal models, such as directed acyclic graphs (DAGs), to reason the cause–effect relationships and design analyses that mitigate confounding bias. Identification of covariate adjustment sets is the most common application of DAG models [1]. Currently, most epidemiologists build DAGs manually, relying on their expert knowledge. But even the most qualified experts may struggle to produce an accurate description of a complex causal process that can involve dozens of variables and hundreds of edges. Indeed, this daunting task leads some researchers to dismiss DAG-based causal inference altogether [2]. ‘Causal structure learning’ is an approach that promises to partly identify DAGs or similar models directly from data. The first causal structure-learning algorithms were proposed >30 years ago [3]; new methods for structure learning continue to appear every year, primarily in the machine-learning literature, and mature, user-friendly software packages such as TETRAD [4], bnlearn [5], and pcalg [6] make using these methods relatively easy. Despite all this, applications of these methods in epidemiology remain rare (compared with other fields such as computational biology) and this could be partly because current structure-learning algorithms have important limitations that hinder their application to complex population data. The paper by Andrews et al. [7] is a good example of current efforts to overcome these limitations and make structure learning more practical.