An efficient private FIM on hadoop MapReduce
Trupti Vishwambhar Kenekar, A.R. Dani · 2016
The frequent item sets can provide valuable business and economic insights. It can be also be useful for research purposes. Privacy is also important for rapidly growing of unstructured non-numeric data which is referred as Big Data. In mining this data objective is to handle such data sets efficiently and preserve privacy in case such large dataset contains sensitive information. Differential privacy is used to protect sensitive information of individual when data released. Proposed system uses non-numeric dataset. Proposed system performs initial processing on non-numeric dataset and finds frequent itemset. Work is done for applicability of FIM techniques on MapReduce platform. Map Reduce is used to implement the parallelization of FP-growth algorithm, thereby improving the overall performance of FIM. The experimental results show that our proposed system is feasible and valid with good speedup and higher mining efficiency, and can meet the rapidly growing needs of frequent itemset mining for massive small files dataset.