A New Fuzzy Clustering-Based Imputation Method
Thanh Le, Lan Vu · 2018
Fuzzy clustering has been used in numerous research disciplines and commercial applications to identify groups of real-world objects. Most fuzzy clustering algorithms require complete datasets; however, real-world datasets may have missing values due to technical limitations. To address this problem, we present a new algorithm where data are clustered using the Fuzzy C-Means algorithm, followed by approximating the fuzzy partition by a probabilistic data distribution model which is then used for missing value imputation as well as for defuzzification. Using distribution-based approach, our method is most appropriate for datasets where the data are non-uniform. We show that our method outperforms seven popular imputation algorithms on uniform and non-uniform artificial datasets as well as real datasets with unknown data distribution model.