Fuzzy data mining: effect of fuzzy discretization
Hisao Ishibuchi, Takashi Yamamoto, Tomoharu Nakashima · 2002
When we generate association rules, continuous attributes have to be discretized into intervals while our knowledge representation is not always based on such discretization. For example, we usually use some linguistic terms (e.g., young, middle age, and old) for dividing our ages into some fuzzy categories. We describe the extraction of linguistic association rules and examine the performance of extracted rules. First we modify the definitions of the two basic measures (i.e., confidence and support) of association rules for extracting linguistic association rules. The main difference between standard and linguistic association rules is the discretization of continuous attributes. We divide the domain interval of each attribute into some fuzzy regions (i.e., linguistic terms) when we extract linguistic association rules. Next, we compare fuzzy discretization with standard non-fuzzy discretization through computer simulations on a pattern classification problem with many continuous attributes. The classification performance of extracted rules on unseen test patterns is examined under various conditions. Simulation results show that linguistic association rules with rule weights have high generalization ability even when the domain of each continuous attribute is homogeneously partitioned.