Complexity of Rule Sets Induced from Incomplete Data with Attribute-concept Values and "Do Not Care" Conditions
Patrick G. Clark, Jerzy W. Grzymala‐Busse · 2014
In this paper we study incomplete data sets with two interpretations of missing attribute values: attribute-concept values and ``do not care'' conditions. As follows from the recent research, performance in terms of the error rate, for both interpretations, is not significantly different. Hence we decided to study the complexity of rule sets induced from incomplete data sets with these two interpretations of missing attribute values. Experiments were conducted on 176 data sets, using three kinds of probabilistic approximations (lower, middle and upper) and the MLEM2 rule induction system. In our experiments, the size of the rule set was smaller for attribute-concept values for 12 combinations of the type of data set and approximation, for one combination the size of the rule sets was smaller for ``do not care'' conditions and for the remaining 11 combinations the difference in performance was statistically insignificant (5\% significance level). The total number of conditions was smaller for attribute-concept values for ten combinations, for two combinations the total number of conditions was smaller for ``do not care'' conditions, while for the remaining 12 combinations the difference in performance was statistically insignificant. Thus, we may claim that attribute-concept values are better than ``do not care'' conditions in terms of rule complexity.