Determining the appropriate number of nodes for fast mining of frequent patterns in distributed computing environments

Wei-Tee Lin, Chih‐Ping Chu · International Journal of Parallel Emergent and Distributed Systems · 2014

The rapid growth of data brought new challenges, the execution efficiency and scalability, for mining frequent patterns (FPs). To accelerate the execution, many algorithms based on the generate-and-test approach or FP-growth utilising parallel distributed technologies have been proposed in recent years. Most of the past studies focused on designing efficient mining algorithm, and none of them has explored how the appropriate number of computing nodes is determined. Using too many computing nodes will increase the execution time because the existing algorithms need to transmit the sub-databases or FP-trees over the network; using insufficient computing nodes may not effectively distribute the mining loading. In this article, we propose a novel algorithm for efficiently mining FPs with the ability to determine the appropriate number of computing nodes in distributed computing environments. Through empirical evaluations in various simulation conditions, the proposed algorithm is shown to deliver excellent performance in terms of execution time.

Read the paper · More papers on PaperTik