Approximate frequent itemsets mining on data streams using hashing and lexicographie order in hardware

Lázaro Bustio-Martínez, René Cumplido, Martín Letras, Claudia Feregrino Uribe, Raudel Hernández-León, José M. Bande-Serrano · 2017

Frequent Itemsets Mining is a Data Mining technique that has been employed to extract useful knowledge from datasets; and recently, from data streams. Data streams are an unbounded and infinite flow of data arriving at high rates; therefore, traditional Data Mining approaches for Frequent Itemsets Mining cannot be used straightforwardly. Finding alternatives to improve the discovery of frequent itemsets on data streams is an active research topic. This paper introduces the first hardware-based algorithm for such task. It uses the top-k frequent 1-itemsets detection, hashing and the lexicographic order of received items. Experimental results demonstrate the viability of the proposed method.

Read the paper · More papers on PaperTik