Tree-based algorithms for missing data imputation
María Jesús Bárcena Ruiz, Fernando Tusell · COMPSTAT · 2000
Let X be a N × ( p+q ) data matrix, with entries partly missing in the last q columns. A problem of practical relevance is that of drawing inferences from such an incomplete data set. We propose to use a sequence of trees to impute missing values. Essentially, the two algorithms we introduce can be viewed as predictive matching methods. Among their advantages, is their flexibility, which makes no assumptions about the type or distribution of the variables.