FUZZY ASSOCIATION RULE REDUCTION USING CLUSTERING IN SOM NEURAL NETWORK

Marjan Kaedi, Mohammad Ali Nematbakhsh, Nasser Ghasem-Aghaee · 2008

ABSTRACT The major drawback of fuzzy data mining is that after applying fuzzy data mining on the quantitative data, the number of extracted fuzzy association rules is very huge. When many association rules are obtained, the usefulness of them will be reduced. In this paper, we introduce an approach to reduce and summarize the extracted fuzzy association rules after fuzzy data mining. In our approach, in first, we encode each obtained fuzzy association rule to a string of numbers. Then we use self-organizing map (SOM) neural network iteratively in a tree structure for clustering these encoded rules and summarizing them to a smaller collection of fuzzy association rules. This approach has been applied on a data base containing information about 5000 employees and has shown good results. KEYWORDS Fuzzy Data Mining, Rule Reduction, SOM. 1. INTRODUCTION Data mining is a process of discovering various models, summaries, and derived values from a given collection of data (Glenn 2006). Data mining techniques have been developed to turn data into useful task-oriented knowledge. Associations reflect relationships among items in databases, and have been widely studied in the fields of knowledge discovery and data mining. Most algorithms for mining association rules identify relationships among transactions using binary values. Transactions with quantitative values and items are, however, commonly seen in real-world applications. Recent years have witnessed many efforts on discovering fuzzy associations, aimed at coping with fuzziness in knowledge representation and decision support processes (Delgado et al. 2005). Fuzzy association rules described by the natural language are well suitable for the thinking of human subjects. According to (Delgado et al. 2005), (Lee and Kwang 1997) is the first paper introducing fuzzy sets into association rules to diminish the granularity of quantitative attributes. The model uses a membership threshold to change fuzzy transactions into crisp ones before looking for ordinary association rules in the set of crisp transactions. Items keep being pairs, i.e. attribute, label. In (Hong et al.1999), only one item per attribute is considered: the pair (attribute, label) with greater support among those items based on the same attribute. The model is the usual generalization of support and confidence based on sigma-counts. The proposed mining algorithm first transforms each quantitative value into fuzzy sets in linguistic terms and then calculates the scalar cardinalities of all linguistic terms in the transaction data. (Au and Chan 1999) presented a novel algorithm, called FARM, which employs linguistic terms to represent the revealed regularities and exceptions. FARM employs adjusted difference analysis to identify interesting associations among attributes without using any user supplied thresholds. In (Hong et al. 2003), a fuzzy multiple-level data-mining algorithm was proposed that can process transaction data with quantitative values and discover interesting patterns among them. If mining procedure also produces a huge number of rules, a human user does not have the ability to analyze these rules. However, if such a huge number of rules do exist in the data, it will not be appropriate to arbitrarily discard any of them or to generate only a small subset of them. It is much more desirable if we can summarize them. In (Farzanyar et al. 2006), the size of average transactions and original dataset is reduced by the recognition and fusion of similar behaving attributes and mining is performed on the reduced dataset that produces a much smaller but richer set of fuzzy association rules which has been approved by

Read the paper · More papers on PaperTik