Association Rule Mining using Apriori for Large and Growing Datasets under Hadoop
Aruna Govada, Abhinav Patluri, Atmika Honnalgere · 2017
The time consumed by Apriori algorithm to process transactional databases for frequent item set mining grows exponentially with the dimensions in the database. The standard rule mining procedure based on Apriori is suboptimal with respect to large datasets which grow periodically with addition of new data. Our approach to optimization introduces a merging function that merges the existing frequent itemsets with the itemsets generated from the newly obtained data such that any changes in trends are reflected. The proposed procedure can reduce the time consumption of Apriori based association rule mining by up to 50% while still maintaining substantial similarity in output, contingent upon user input. Furthermore, the procedure remains independent of the mechanism of Apriori algorithm; the algorithm may be replaced by any other rule mining algorithm without altering the procedure itself.