A Dual Learning Framework for Multi-Class Imbalanced Datasets with Limited Labeled Data
M. S. Neethu, Vinod Chandra S S · 2025
The increasing availability of unlabeled data in real-world scenarios creates challenges for traditional supervised learning methods that rely heavily on labeled datasets. This work proposes a hybrid framework combining semi-supervised learning and supervised learning techniques to address these challenges, particularly for multi-class classification problems with labeled data scarcity and class imbalance. This study integrates label propagation technique to generate pseudo-labels for unlabeled data and employs supervised classifier models, including XGBoost, random forest, and a stacked ensemble model, for classification tasks. Evaluation across diverse multi-class datasets from the Knowledge Extraction based on Evolutionary Learning (KEEL) data repository demonstrates that the proposed approach achieves substantial performance improvements, with accuracy gain of 8% through 16.35% compared to existing methods and best performance in certain scenarios. These results highlight the ability of framework to optimize the use of unlabeled data and address practical challenges. Its modular design enables integration with any supervised learning approach, providing a flexible and scalable solution for various applications, which could be considered for future work.