Parallel mining algorithms for generalized association rules with classification hierarchy

Takahiko Shintani, Masaru Kitsuregawa · 1998

Association rule mining recently attracted strong attention. Usually, the classification hierarchy over the data items is available. Users are interested in generalized association rules that span different levels of the hierarchy, since some-times more interesting rules can be derived by taking the hierarchy into account. In this paper, we propose the new parallel algorithms for mining association rules with classification hierarchy on a shared-nothing parallel machine to improve its performance. Our algorithms partition the candidate itemsets over the processors, which exploits the aggregate memory of the sys-tem effectively. If the candidate itemsets are partitioned without considering classification hierarchy, both the items and its all the ancestor items have to be transmitted, that causes prohibitively large amount of communications. Our method minimizes interprocessor communication by consid-ering the hierarchy. Moreover, in our algorithm, the avail-able memory space is fully utilized by identifying the fre-quently occurring candidate itemsets and copying them over all the processors, through which frequent itemsets can be processed locally without any communication. Thus it can effectively reduce the load skew among the processors. Sev-eral experiments are done by changing the granule of copying itemsets, from the whole tree, to the small group of the fre-quent itemsets along the hierarchy. The coarser the grain, the easier the control but it is rather difficult to achieve the sufficient load balance. The finer the grain, the more com-plicated the control is required but it can balance the load quite well. We implemented proposed algorithms on IBM SP-2. Per-formance evaluations show that our algorithms are effective for handling skew and attain sufficient speedup ratio. 1

Read the paper · More papers on PaperTik