Mining High-Utility Sequential Patterns from Big Datasets

Jerry Chun‐Wei Lin, Yuanfa Li, Philippe Fournier‐Viger, Youcef Djenouri, Leon Shyue-Liang Wang · 2019

High-Utility Sequential Pattern Mining (HUSPM) has become an emerging issue in recent decades since it reveals more information such as the utility and sequence factors for knowledge discovery. For the previous works, many algorithms were presented to speed up the mining performance regarding a single machine with small datasets. In real-world applications, the size of dataset can be collected from many places or devices, such as PC, Internet of Things (IoT), mobile devices, and shopping malls, among others. It is necessary to build an efficient model to handle the big dataset for HUSPM. In this paper, we present a four-stages MapReduce framework based on the Spark platform for mining the high-utility sequential patterns from a very large database. From the experimental results, we then can observe that the designed model outperforms the state-of-the-art approaches for handling the very big dataset.

Read the paper · More papers on PaperTik