Implementation of an improved algorithm for frequent itemset mining using Hadoop
Ruchi Agarwal, Sunny Singh, Satvik Vats · 2016
Searching frequent item-sets in large size heterogeneous databases in minimal time is considered as one of the most important data mining problem. As a solution of this problem, various algorithms have been proposed to speed up execution. Most of the recent proposed algorithms focussed on parallelizing the workload using large number of machine in distributed computational environment like MapReduce framework. A few of them are actually capable to determine the appropriate number of required computing computers, considering workload balancing and execution efficiency. But internally not capable to determine exact number of required iteration for any large size datasets in advance to find out the frequent item-set based on iterative sampling. In this paper, we propose an improved and compact algorithm (ICA) for finding frequent item-set in minimal time, using distributed computational environment. It is also capable of determining the exact number of internal iteration required for any large size datasets whether data is in structured or unstructured format.