New Approach in Big Data Mining for Frequent Itemset Using Mapreduce in HDFS
Pallavi V. Nikam, Deepa S. Deshpande · 2018
To find out the frequent itemset is very important task in data mining. These frequent itemsets are useful in applications like Association rule mining and co-relations. To extract frequent itemsets these systems are using some algorithms, but these are inefficient in distributing and balancing the load, when it comes across excessive data. Automatic parallelization is also not possible with these algorithms. To overcome these issues of existing algorithms there is need to develop algorithm which will support the missing features, such as automatically parallelization, balancing and good distribution of data. In this paper we are using a new approach to find out frequent itemsets by using MapReduce. Modified Apriori algorithm is used with HDFS environment this is called FiDoop Technique. In this technique mapreduce process will work independently and concurrently by using the decompose strategy. The result of this mapreduce technique will be given to the reducers and reducers will show the result. In the experiment we used three different algorithms like basic apriori, FP Growth and our proposed modifies apriori, the system has executed in standalone machine as well as distributed environment and shown the results how proposed algorithm is better than existing algorithms.