Fast algorithms for mining association rules and sequential patterns

Ramakrishnan Srikant, Jeffrey F. Naughton · 1996

This dissertation presents fast algorithms for mining associations in large datasets. An example of an association rule may be \\30 % of customers who buy jackets and gloves also buy hiking boots. " The problem is to nd all such rules whose frequency is greater than some user-speci ed minimum. We rst present a new algorithm, Apriori, for mining associations between boolean attributes (called items). Empirical evaluation shows that Apriori is typically 3 to 10 times faster than previous algorithms, with the performance gap increasing with the problem size. In many domains, taxonomies (isa hierarchies) on the items are common. For example, a taxonomy maysaythat jackets isa outerwear isa clothes. We extend the Apriori algorithm to nd associations between items at any level of the taxonomy. Next, we consider associations between quantitative and categorical attributes, not just boolean attributes. We deal with quantitative attributes by

Read the paper · More papers on PaperTik