Imputation of missing healthcare data

Mohaimanul Hoque Chowdhury, M. Kamrul Islam, Shahidul Islam Khan · 2017

In the field of data mining, missing values has always been a crucial factor. Incorrect imputation of missing values could lead to inaccurate research as well as wrong predictions. Developing a generalized imputation strategy that can be used across a variety of dataset is very hard. Because each dataset has its attributes and characteristics, and finding another dataset of similar property could be a very hard task. In Bangladesh, real healthcare data is very noisy which makes the knowledge discovery by health researchers very difficult. One of the main sources of noise in healthcare data is missing values. In this paper, we presented a general model for the imputation of any kind of missing data. We also implemented three different algorithms namely Amelia, FURIA, and MICE in real healthcare dataset to impute missing values. Experimental results on 65000 patient records show that MICE algorithm performs better among the three.

Read the paper · More papers on PaperTik