A Hybrid Labeling Strategy for Imbalanced Data Stream in Presence of Concept Drifts
Bohnishikha Halder, K. M. Azharul Hasan · 2022
Imbalanced data stream with concept drift is a tough challenge to predict correct labels. Handling both of these issues with limited labeling budgets gets more attention nowadays. Though some existing methods try to address these problems at the same time, it fails to provide appropriate outcomes for imbalanced data with a lower budget. So, to get higher performances with lower labeling budgets we propose a method is proposed. A cluster-based initialization technique is considered to train an ensemble classifier with the most representative instances. This cluster-based initialization process helps to be robust against the concept drift and reduce the overall labeling cost in active learning situation. In addition, a noble imbalance strategy is introduced that provides a special priority to the minority class data. For measuring the efficiency of the method, four synthetic datasets generator with abrupt and gradual drifting and two real-world datasets are considered. The proposed method attains higher AUC values for all the datasets with the least labeling cost. And, every dataset shows better recall values as compare to the other related methods.