A parallel algorithm of association rules based on cloud computing
Wang Yong, Zhe Zhang, Wang Fang · 2013
In view of the traditional parallel FP-growth algorithm (PFP)that suffers from two major limitations, namely, multiple database scans requirement (i.e., high I/O cost) and high inter-processor communications cost, therefore we design and implement a parallel association rules mining method based on cloud computing. The algorithm adopts the separation strategy to simply visit a local database only once, thus, the inter-processor communication I/O overhead is reduced. What's more, the MapReduce model is used to solve the problem of huge amounts of data mining, as well as the calculated execution taking place in the local data storage node, which can avoid large amounts of data on the network transmission and reduce the communication overhead. By using ordinary PC structures, Hadoop cluster experimental results verify that the proposed algorithm based on cloud computing offers higher efficiency and has a good speedup.