Mining of High-Utility Sequence Patterns in Large-Scale Uncertain Databases

Jimmy Ming‐Tai Wu, Shuo Liu, Jerry Chun‐Wei Lin · 2022 IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf on Pervasive Intelligence and Computing, Intl Conf on Cloud and Big Data Computing, Intl Conf on Cyber Science and Technology Congress (DASC/PiCom/CBDCom/CyberSciTech) · 2022

In the age of the Internet of Things (IoT), information is collected from sensor devices, which can lead to data loss or data in doubt, among other things. We need to use probability concept to accurately describe the uncertain data we have collected so that we can successfully search a vast undetermined data warehouse for information that can be used in production and applications. As the data in the database may be arranged as a particular order/sequence, whether temporally or spatially, in the field of data processing, the High Utility-Probability Sequential Pattern Mining (HUPSPM) has become an important new research topic and has been widely studied in recent years. The timestamps has led to the development of a considerable number of efficient algorithms for sequential mining. However, these methods have several drawbacks, and most of them cannot handle the very huge datasets. For this reason, the implementation of an advanced MapReduce framework for big data processing is the solution to the problems caused by the currently used approaches. The method presented in this paper has the potential to eliminate the need for multiple database scans, to divide the database into several partitions, and to extend the capabilities of parallel computing. The original database is "pruned" according to the previously defined method that can be used to reduce the total number of potential candidates in a time and resource efficient manner. The approach presented in this paper has been shown in experiments to be very beneficial in finding sequences with high utility and probability in huge datasets.

Read the paper · More papers on PaperTik