Tree-based algorithms for missing data imputation

María Jesús Bárcena Ruiz, Fernando Tusell · COMPSTAT · 2000

Let X be a N × ( p+q ) data matrix, with entries partly missing in the last q columns. A problem of practical relevance is that of drawing inferences from such an incomplete data set. We propose to use a sequence of trees to impute missing values. Essentially, the two algorithms we introduce can be viewed as predictive matching methods. Among their advantages, is their flexibility, which makes no assumptions about the type or distribution of the variables.

Read the paper · More papers on PaperTik