Coping With Missing Attribute Values Based on Closest Fit in Preterm Birth Data: A Rough Set Approach
Jerzy W. Grzymala‐Busse, Witold J. Grzymała-Busse, Linda Goodwin · Computational Intelligence · 2001
Data mining is frequently applied to data sets with missing attribute values. A new approach to missing attribute values, called closest fit, is introduced in this paper. In this approach, for a given case (example) with a missing attribute value we search for another case that is as similar as possible to the given case. Cases can be considered as vectors of attribute values. The search is for the case that has as many as possible identical attribute values for symbolic attributes, or as the smallest possible value differences for numerical attributes. There are two possible ways to conduct a search: within the same class (concept) as the case with the missing attribute values, or for the entire set of all cases. For comparison, we also experimented with another approach to missing attribute values, where the missing values are replaced by the most common value of the attribute for symbolic attributes or by the average value for numerical attributes. All algorithms were implemented in the system OOMIS. Our experiments were performed on the preterm birth data sets provided by the Duke University Medical Center.