Distributed and parallel high utility sequential pattern mining
Morteza Zihayat, Zane Zhenhua Hut, Aijun An, Yonggang Hut · 2016
The problem of mining high utility sequential patterns (HUSP) has been studied recently. Existing solutions are mostly memory-based, which assume that data can fit into the main memory of a computer. However, with advent of big data, such an assumption does not hold any longer. Hence, existing algorithms are not applicable to the big data environments, where data are often distributed and too large to be dealt with by a single machine. In this paper, we propose a new framework for mining HUSPs in big data. A distributed and parallel algorithm called BigHUSP is proposed to discover HUSPs efficiently. At its heart, BigHUSP uses multiple MapReduce-like steps to process data in parallel. We also propose a number of pruning strategies to minimize search space in a distributed environment, and thus decrease computational and communication costs, while still maintaining correctness. Our experiments with real life and large synthetic datasets validate the effectiveness of BigHUSP for mining HUSPs from large sequence datasets.