Discovering linear causal model from incomplete data

Gang Li, Honghua Dai, Yahan Tu · Deakin Research Online (Deakin University) · 2003

\t\t\t\t\tOne common drawback in algorithms for learning Linear Causal Models is that they can not deal with incomplete data set. This is unfortunate since many real problems involve missing data or even hidden variable. In this paper, based on multiple imputation, we propose a three-step process to learn linear causal models from incomplete data set. Experimental results indicate that this algorithm is better than the single imputation method (EM algorithm) and the simple list deletion method, and for lower missing rate, this algorithm can even find models better than the results from the greedy learning algorithm MLGS working in a complete data set. In addition, the method is amenable to parallel or distributed processing, which is an important characteristic for data mining in large data sets. \t\t\t\t

Read the paper · More papers on PaperTik