Frequent and non-frequent pattern detection in big data streams: An experimental simulation in 1 trillion data points

Konstantinos F. Xylogiannopoulos, Reda Alhajj, Panagiotis Karampelas · 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM) · 2016

Big data streaming analysis nowadays has become one of the most important topic in the list of data analysts since enormous amount of data are produced daily by the numerous smart devices. The analysis of such data is very important and the detection of frequent or even non-frequent patterns can be critical for many aspects of our lives. In the current paper, we propose a new methodology based on our previous work regarding the detection of all repeated patterns in a string in order to analyze a very big data stream with 1 Trillion digits, composed from 1 thousand subsequences of 1 billion digits each one. More specifically, using the novel data structure, LERP Reduced Suffix Array, and the innovative ARPaD algorithm which allows the detection of all repeated patterns in a string we managed to analyze each one of the 1 billion data points, using 10 computers with standard hardware configuration, in 33 minutes which outperforms to the best of our knowledge any other existing methodology, which is equivalent to data point generation every 2 microseconds.

Read the paper · More papers on PaperTik