Weighted Frequent Pattern Mining using RDD, the Basic Spark Abstraction

Ambily Mohan, R.L. Visakh · 2014

Data mining [1] is the compilation of techniques for the creative, automatic discovery of previously unknown, fitting, new, and understandable patterns that are present in large databases. Frequent pattern mining is considered as an important task in data mining. Patterns that occur frequently together in a data set can be regarded as a frequent pattern. Traditional frequent pattern mining algorithms treated patterns and items within the patterns the same with the assumption that all items have same importance. However this is not true in the case of real world data, as there are items with different importance based on the profit they incur. This resulted in assigning weights to the items to reflect their importance in real world. This work focuses on how the size of memory can be minimized while extracting frequent weighted itemsets from weighted item transaction databases. The algorithm makes use of vertical data format where each item is stored along with the ids of all transactions (or tids) containing the item. When the tidset cardinality gets very large the size of memory required to store intermediate results becomes very large. To tackle this problem, a new approach called Resilient Distributed Dataset (RDD), the basic abstraction in Spark is used.

Read the paper · More papers on PaperTik