Improvement parallelization in Apriori Algorithm
Qiu Hui-qi · 2020
When people look for the internal relationship of massive data, the classical Apriori association algorithm can not meet the mining scenario of massive data in terms of algorithm efficiency and I/O performance of data processing process. When mining and analyzing the classical Apriori association algorithm, all transactions are always located in the file system or database, and each iteration needs to scan the transaction database; when mining frequent sets by Apriori algorithm, a large number of project sets will be generated, resulting in excessive I/O overhead and excessive system resources. This paper proposes a distributed apriori algorithm based on MapReduce. The Hadoop distributed file system HDFS is used to automatically realize the distributed storage (fragmentation) of big data. Combined with MapReduce computing framework, map and reduce are used for parallel processing in generating candidate itemsets and computing frequent itemsets respectively. The distributed system is fully utilized to improve the processing ability, At the same time, the monotonicity of frequent itemsets is used to optimize the transaction database. Experimental analysis, the addition of compute nodes can significantly provide mining algorithm performance. At the same time, the improved Apriori algorithm has a great improvement in operating efficiency when it conducts association rule analysis on the data set with a large amount of item indexes.