Multi-classifier semi-supervised data stream classification algorithm based on online learning
Chaofan Liu, Xianning Lin · 2024
Massive data is continuously generated from a variety of real-world scenarios and fields in all walks of life, including data recorded in hardware sensors, huge amounts of text, voice and image data, various user data generated by commonly used systems and software. The common denominator of these data is continuous generation, large volume, and lack of correct and reliable class labels. In the research progress of data stream classification algorithms, there are relatively mature solutions to solve the problem of rapid data generation and large volume, but how to accurately classify data instances when class labels are delayed or directly missing is still a problem worth studying. To handle these challenges, a multi-classifier semi-supervised data stream classification algorithm based on continuous learning is proposed. The classifier adopts an ensemble pool framework and uses the principle of majority voting to label the newly arrived instances. After the instances get a small number of labels, the current base classifier is updated by calculating the forgetting measurement index, and a large number of unlabeled instances are utilized to participate in the continuous learning process of the base classifier. Evaluations on benchmark datasets shows the advantages of the algorithm. In the case of using a relatively simple ensemble model, a simple update strategy and implicit concept drift adaptation are adopted, and the classification accuracy of the algorithm on the data stream can be ahead of the existing algorithm.