On ensemble components selection in data streams scenario with reoccurring concept-drift
Piotr Duda, Maciej Jaworski, Leszek Rutkowski · 2017
In this article we consider the problem of data stream classification with recurring concept-drift. The proposed method determines when to add or remove a component from an ensemble. The algorithm discussed in this article is an extension of the ASE (Automatically Sized Ensemble) algorithm which guarantees that, with the assumed probability, adding a new component increases the accuracy not only for the current data chunk, but also for the entire data stream. The new approach improves the adaptation of the ASE algorithm to non-stationary environments, particularly when the data stream is generated by different data distributions that appear and disappear alternately. The procedure uses a Kullback-Leibler discrepancy to measure the suitability of elements for the considered data distribution.