Supervised knowledge discovery from incomplete data

Alexandros Kalousis, Mélanie Hilario · International Conference on Data Mining · 2000

Incomplete data can raise more or less serious problems in knowledge discovery systems depending on the quantity and pattern of missing values as well as the generalization method used. For instance, some methods are inherently resilient to missing values while others have built-in methods for coping with them. Still others require that none of the values are missing; for such methods, preliminary imputation of missing values is indispensable. After a quick overview of current practice in the machine learning field, we explore the problem of missing values from a statistical perspective. In particular, we adopt the well-known distinction between three patterns of missing values-missing completely at random (MCAR), missing at random (MAR) and not missing at random (NMAR)-to focus a comparative study of eight learning algorithms from the point of view of their tolerance to incomplete data. Experiments on 47 datasets reveal a rough ranking from the most resilient (e.g., Naive Bayes) to the most sensitive (e.g., IB1 and-surprisingly-C50rules). More importantly, results show that for a given amount of missing values, their dispersion among the predictive variables is at least as important as the pattern of missingness.

Read the paper · More papers on PaperTik