A Fuzzy C-Means Approach for Incomplete Data Sets Based on Nearest-Neighbor Intervals

Dan Li, Chong Quan Zhong, Shi Qiang Wang · Applied Mechanics and Materials · 2013

Partially missing data sets are a prevailing problem in pattern recognition. In this paper, the problem of clustering incomplete data sets is considered, and missing attribute values are imputed by the centers of corresponding nearest-neighbor intervals. Firstly, the algorithm estimates the nearest-neighbor intervals of missing attribute values by using the attribute distribution information of the data sets sufficiently. Secondly, the missing attribute values are imputed by the center of the intervals so as to clustering incomplete data sets. The proposed algorithm introduces the nearest neighbor information into incomplete data clustering, and the comparisons of the experimental results for two UCI data sets demonstrate the capability of the proposed algorithm.

Read the paper · More papers on PaperTik