Evolving clustering based data imputation

Chandan Gautam, Vadlamani Ravi · 2014

Missing data is an inevitable problem in many disciplines. In this paper, we employed an Evolving Clustering Method (ECM) based imputation method and performed sensitivity analysis of the influence of threshold value (Dthr) on imputation results over 12 datasets. We experimented on a large range of Dthr values from 0.001 to 0.999, in steps of 0.001, in order to see which value of Dthr would perform better imputation compared to K-Means+MLP. Thereby, we provided an upper bound for the Dthr value in ECM algorithm. Further, we tested the effectiveness of the online clustering based imputation method on 12 datasets under 10-fold cross validation set up. ECM yielded better performance compared to K-Means + Multilayer perceptron hybrid algorithm, appearing in literature. It is due to strong local learning capability of ECM and selection of an optimal Dthr value.

Read the paper · More papers on PaperTik