Statistical and Probabilistic Methods for Data Stream Mining
Katharina Tschumitschew · Digitale Bibliothek Braunschweig (Verbundzentrale Göttingen (VZG)) · 2012
The aim of this work is not only to highlight and summarize issues and challenges which arose during the mining of data streams, but also to find possible solutions to illustrated problems. Due to the streaming nature of the data, it is impossible to hold the whole data set in the main memory, i.e. efficient on-line computations are needed. For instance incremental calculations could be used in order to avoid to start the computation process from scratch each time new data arrive and to save memory. Another important aspect in data stream analysis is that the data generating process does not remain static, i.e.\ the underlying probabilistic model cannot be assumed to be stationary. The changes in the data structure may occur over time. Dealing with non-stationary data requires change detection and on-line adaptation. Furthermore real data is often contaminated with noise, this causes a specific problem for approaches dealing with the data streams. They must be able to distinguish between changes according to noise and changes of the underlying data generating process or its parameters. In this work we propose a variety of different methods, which fulfil specific requirements of data stream mining. Furthermore we carry out theoretical analysis of effects of noise and changes in data stream for sliding window based evolving system in order to illustrate the problem of suboptimal window size. In order to do the validation of an evolving system significant, we propose some simple benchmark tests that can give an idea of how much an evolving system might be misled by noise.