Analysis of Sequential Pattern Mining Algorithms

Alpa Reshamwala, Neha Mishra · 2014

Abstract- Sequential pattern mining is an important data mining problem with broad applications. Most of the previously developed sequential pattern mining methods, such as SPAM and SPADE, explore a candidate generation-and-test approach [12] which reduces the number of candidates to be examined. In this paper, we have implemented SPADE, SPAM and Prefixspan algorithm on the two databases. One database is sign database which is taken from ASL (American sign language database) [11]. The second dataset is Kosarak dataset containing 10000 sequences of click-stream data from an hungarian news portal. Sign dataset forms the dense dataset with few distinct items and Kosarak forms the sparse dataset with maximum distinct items. From the experimental results, SPADE performs better in both the dense as well as sparse dataset taken for simulation study. Performance of SPAM is worst when executed on sparse dataset. The number of sequences generated is same in both the dataset by all the mentioned algorithms. For dense dataset prefixsapn uses less memory whereas in sparse dataset it utilizes the most. In Dense dataset SPAM and SPADE are utilizing approximately constant memory. In sparse dataset minimum utilization of memory is by SPADE.

Read the paper · More papers on PaperTik