Finding Association Rules from Quantitative Data Using Data Booleanization

Susan P. Imberman, Bernard Domanski · Journal of the Association for Information Systems · 2001

Finding association rules in data that is naturally binary has been well researched and documented. Finding association rules in numeric/categorical data has not been as easy. Many quantitative algorithms work directly on the numeric data limiting the complexity of the generated rules. In addition, as you create intervals from the numeric data the dimensionality of the problem increases significantly, causing execution time to blow up. Quantitative data can be "booleanized" using simple thresholds as the basis for boolean classification. The "booleanized" data can be used in association rule algorithms to find interesting rules and patterns in this data. Once significant associations are found, we can increase the dimensionality on the selected interesting variables. We use an association rule algorithm, Boolean Analyzer, to look at rules. Finding significant rules is very dependent on how thresholds for booleanization are defined. We investigate the association rules generated by this algorithm when thresholds are defined by experts, and compare these rules to those calculated using mode, mean, median as threshold measures, as well as to rules derived by using thresholds found by k-means clustering. Keywords: data mining, booleanization, association rules, dependency rules 1.

Read the paper · More papers on PaperTik