Frequent items mining on data stream using hash-table and heap
Shan Zhang, Chen Ling, Li Tu · 2009
Most of the existing algorithms for mining frequent items on data stream do not emphasis the importance of the recent data items. We present an algorithm to detect the items with frequency counts exceeding a user-specified threshold. Our algorithm uses a hash table L and a heap to record the potential frequent items, and can detect ¿-approximate frequent data items on data stream using O(|L|+ ¿-1) memory space and the processing time for each data item is O(log¿-1). Experimental results on several artificial and real datasets show our algorithm has higher precision, requires less memory and consumes less computation time than other similar methods.