P-EAARM: A Generic Framework Based on Spark for EAs-based Association Rule Mining

Tao Zhang, Mingli Shi, Jingwei Wang, Guangming Yang · 2019

Many researchers have proposed serialized evolutionary algorithms (EAs) which are utilized to discover association rules from datasets. However, serialized algorithms are ineffective to tackle with large datasets because of the hardware limitations. Aiming at this issue, we propose a generic framework based on Spark for EAs-based association rule mining (P-EAARM). Based on the mechanism of Resilient Distributed Dataset in Spark, large datasets can be scanned in parallel. The basic idea of P-EAARM is to reduce the computational cost when each idea of EAs scans the dataset to calculate measures. Instead of the reduceBy operator of RDD used in the parallel PSO algorithm for association rules mining (PPQAR), P-EAARM uses the mechanism of Accumulator provided by Spark to reduce computational cost. In addition, the final non-dominated solutions are stored as a RDD in order to improve its scalable ability. As a case study, we use a hybrid NSGA-II-MOPSO algorithm and an extended dataset taken from IBM SPSS modeler to validate P-EAARM in terms of computational cost and speedup. The results show that, compared with the scanning method mentioned in PPQAR, P-EAARM reduces the computational cost markedly. Compared with serialized algorithm, P-EAARM can acquire a 430% speedup approximately. The results prove the validity of our framework.

Read the paper · More papers on PaperTik