Missing Values Estimation on Multivariate Dataset : Comparison of Three Type Methods Approach
Yoga Pristyanto, Irfan Pratama · 2019 International Conference on Information and Communications Technology (ICOIACT) · 2019
Knowledge discovery has become ever more essential in the digital era. The very first step of knowledge discovery is data acquisition. It can be gathered from all across the field and automatically made the data type is also various. The result of that data acquisition process is an arguably-huge dataset. Observed data can be achieved by several method, such as censor record or by frequent observation. Each of the analysis process consists of several steps, one of them is preprocessing. Preprocessing is a step or phase to identify, selection, or problem handling of the data. Missing values handling are included in the preprocessing step. The purpose of this research is to find out which type of approach of missing values handling work better on this type of dataset. This research uses three approaches and been compared to each other such as Mode Imputation, Decision tree, and Class Center based Missing Values Imputation. To perform a fair comparison among them, several scenarios of missing values appearance have been made. Dataset scenarios for this study are actually artificially "deleted" to be able to measure the performance of the methods. From the evaluation process, Decision Tree method shows a consistency even on different missing point's amount. Numerically have a slightly lower that the other methods.