Strategic Construction of Initial Datasets for Active Learning: Leveraging Self-Supervised Learning
Sekjin Hwang, Daeyoung Heo, Joonsoo Choi, Joonsoo Choi · IEEE Access · 2026
Abstract Deep learning has demonstrated remarkable achievements across various fields. However, its success heavily relies on the availability of large-scale labeled data. Labeling data is a time-consuming and costly process, prompting numerous studies aimed at reducing these expenses. Active learning is a prominent data-efficient learning methodology that has garnered significant attention. Active learning methods iteratively select data that are most effective for training models, thereby gradually constructing a compact dataset. It typically assumes the presence of a small amount of labeled data at the start of training, and experiments generally use randomly composed initial labeled datasets. Although the importance of initial dataset construction is well recognized because of its impact on the level of model training in most active learning methods, practical research in this area remains limited. In this study, we propose a method of data initialization using self-supervised learning from an active learning perspective. This method focuses on constructing a small initial dataset that maximizes learning efficiency by utilizing an unlabeled dataset. The impact of the proposed method on active learning was evaluated using a representative image classification dataset, which demonstrated significant performance improvements.