Scalable APRIORI-Based Frequent Pattern Discovery

Sean Chester, Ian Sandler, Alex Thomo · 2009

Frequent pattern discovery, the task of finding sets of items that frequently occur together in a dataset, has been at the core of the field of data mining for the past sixteen years. In that time, the size of datasets has grown much faster than has the ability of existing algorithms to handle those datasets. Consequently, improvements are needed. In this paper, we take the classic algorithm for the problem, A priori, and by adding a vertical sort drastically improve its performance characteristics when processing very large datasets. We use the benchmark large dataset webdocs from the FIMI 2004 conference to contrast our performance against several state-of-the-art implementations and demonstrate both equal efficiency with lower memory usage at all support thresholds and also the ability to mine support thresholds as yet unattempted in literature. We also indicate how this work can be extended to achieve yet more impressive results.

Read the paper · More papers on PaperTik