Distributed PrefixSpan algorithm based on MapReduce

Yongqing Wei, Dong Liu, Lin-shan Duan · 2012

For mining sequential patterns on massive data set, the distributed sequential pattern mining algorithm based on MapReduce programming model and PrefixSpan is proposed. Mining tasks are decomposed to many small tasks, the Map function is used to mine each Prefix-Projected sequential pattern, and the projected databases were constructed parallelly. It simplifies the search space and acquires a higher mining efficiency. Then the intermediate values are passed to a Reduce function which merges together all these values to produce a possibly smaller set of values. Both theoretical analyses and experimental results show MR-PrefixSpan reduces the time of scanning database. It solves the problem of mining massive data effectively, has considerable speedup and scaleup performances with an increasing number of processors on the Hadoop platform.

Read the paper · More papers on PaperTik