Discovery of Frequent Pagesets from Weblog Using Hadoop Mapreduce Based Parallel Apriori Algorithm

H K Sowmya, N V Uma Reddy, C. Kavyashree, R J Anandhi · 2022

Web usage mining is a critical stage in analyzing user behavior that involves mining frequently visited pages. Mining usage patterns from a big weblog file using existing frequent mining techniques are inefficient in terms of storage and performance. These algorithms did not perform well while mining big volumes of data due to CPU, storage, and main memory restrictions. To address these challenges, this paper proposes the parallel Apriori technique, which is based on Hadoop MapReduce and is an effective, extensible, and easy programming paradigm for processing large volumes of data. It employs pruning and session reduction techniques to eliminate infrequent pagesets, reducing the time required to process massive amounts of weblog data. It also employs clusters of processing nodes to improve performance. Experiments are conducted to demonstrate performance of the suggested approach using four genuine weblog data sets. Experiments reveal that our suggested system beats existing sequential association rule mining systems in finding similar pagesets, such as Apriori, Eclat, and FP-Growth.

Read the paper · More papers on PaperTik