A Spark-based Incremental Algorithm for Frequent Itemset Mining

Haoxing Wen, Mingdong Kou, Hengyi He, Xiaoguang Li, Huaixiao Tou, Yulu Yang · 2018

Association rule mining plays an important role in many areas, including market basket analysis, intrusion detection, bioinformatics and so on. As an efficient approach of finding frequent itemset among large datasets, several parallel Apriori-based algorithms are widely used in association rule mining. Moreover, datasets are always changed in many real-world applications. For example, datasets of the purchased products from an e-commercial website is growing all the time. However, the existing parallel Apriori-based algorithms cannot update the frequent itemset efficiently for these large and evolving datasets. So we propose an incremental parallel Apriori-based algorithm in this paper. As the datasets increase, our algorithm updates the frequent itemset based on frequent itemset in previous, instead of re-computing the whole datasets from scratch. We implement the proposed algorithm on Spark and evaluate its performance via groups of experiments on some real-world datasets. It is demonstrated by the experimental results that the proposed algorithm improves the performance of mining frequent itemset on the large and evolving data sets significantly.

Read the paper · More papers on PaperTik