Efficient Discovery of Weighted Frequent Itemsets in Very Large Transactional Databases: A Re-visit

Rage Uday Kiran, Amulya Kotni, P. Krishna Reddy, Masashi Toyoda, Subhash Bhalla, Masaru Kitsuregawa · 2018

Weighted Frequent Itemset (WFI) mining is an important model in data mining. The popular adoption and successful industrial application of this model has been hindered by the following two obstacles: (i) finding WFIs is a computationally expensiveness process as these itemsets do not satisfy the downward closure property and (ii) lack of parallel algorithms to find WFIs in very large databases (e.g. astronomical data and twitter data). This paper makes an effort to address these two obstacles. Two pattern-growth algorithms, Sequential Weighted Frequent Pattern-growth and Parallel Weighted Frequent Pattern-growth, have been introduced to discover WFIs efficiently. Both algorithms employ three novel pruning techniques to reduce the computational cost effectively. The first pruning technique prunes some of the uninteresting items by employing a criterion known as cutoff weight. The second pruning technique, called conditional pattern base elimination, eliminates the construction of conditional pattern bases if a suffix item is an uninteresting item. The third pruning technique, called pattern-growth termination, defines a new terminating condition for the pattern-growth technique. Experimental results demonstrate that the proposed algorithms are memory and runtime efficient, and highly scalable as well.

Read the paper · More papers on PaperTik