Chapter 9 Integration of Dataset Scans in Processing Sets of Frequent Itemset Queries

Marek Wojciechowski, Maciej Zakrzewicz, Paweł Boiński · 2012

Frequent itemset mining is often regarded as advanced querying where a user specifies the source dataset and pattern constraints using a given constraint model. In this chapter we address the problem of processing sets of frequent itemset queries, which brings the ideas of multiple-query optimization to the domain of data mining. The most attractive method of solving the prob- lem with respect to possible practical applications is Common Counting which consists in concurrent execution of the queries using Apriori with the integra- tion of scans of the parts of the database shared among the queries. The major advantage of Common Counting over its alternatives is its applicability to arbi- trarily large batches of queries. If the memory structures of all the queries to be processed by Common Counting do not fit together in main memory, the set of queries has to be partitioned into subsets processed in several phases. We for- malize the problem of dividing the set of queries for Common Counting as a specific case of hypergraph partitioning and provide a comprehensive overview of query set partitioning algorithms proposed so far.

Read the paper · More papers on PaperTik