Incorporating an EM-approach for handling missing attribute-values in decision tree induction

Amitava Karmaker, Stephen S. Kwek · 2005

Data with missing attribute-values are quite common in many classification problems. In this paper, we incorporate an expectation-maximization (EM) inspired approach for filling up missing values to decision tree learning with the objective of improving classification accuracy. Here, each missing attribute-value is iteratively filled using a predictor constructed from the known values and predicted values of the missing attribute-values from the previous iteration. We show that our approach significantly outperforms some standard machine learning methods for handling missing values in classification tasks.

Read the paper · More papers on PaperTik