Hierarchical Cluster-Based Adaptive Model for Semi-Supervised Classification of Data Stream with Concept Drift

Keke Qin, Yixiu Qin · 2019

Compared with the research of data stream with concept drift in supervised environment, the work that in semi-supervised environment is more challenging. There is currently very little work in this area, although more meaningful. Considering existing chunk-based processing algorithms are only suitable for periodic concept drift and show poor performance in complex concept drift scenarios, such as concept drift may occur at any time, the duration of each concept is not exactly the same and multiple different types of concept drift types may alternate or mixed at the time. In this work, we propose an online plus offline memory model for processing streaming data in online learning mode. Compared with chunk-based incrementally processing mode, our algorithm does not have to set-up chunk size and model is maintained online. Concept drift processing mechanism is triggered at interval to extract knowledge and clean samples of local area. Therefore, it has more adaptability to complex concept drift scenarios. Extensive experiments on benchmark artificial and real-world datasets shows that our algorithm achieves significantly higher or at least comparable accuracy and behaves more stable in comparison with state-of-the-art cluster-based techniques.

Read the paper · More papers on PaperTik