Active learning evaluation metrics for classification and regression frameworks

Alaa Tharwat, Bjarne Jaster, Wolfram Schenck, Martin Kohlhase · Engineering Applications of Artificial Intelligence · 2026

In recent years, the rapid growth of Internet of Things (IoT) devices, the popularity of social media platforms, and user interactions in various online environments, such as streaming platforms and mobile applications, have led to a significant increase in the amount of data generated. However, as this data is raw and unlabeled, its value for training supervised machine learning models is still limited. The challenge is further complicated by the costly and time-consuming process of manual labeling. Active learning (AL) provides a solution by selectively labeling small but highly informative and representative subset of data points to enable better generalization and improve model performance on unseen data for both classification and regression. The use of AL techniques spans both classification and regression frameworks, each facing different challenges such as binary or multi-class settings, balanced and imbalanced datasets, and the presence of outliers. Such diversity necessitates fair and representative metrics tailored to each AL scenario. Despite this need, some evaluation metrics remain underexplored, and many studies consider only limited perspectives, often evaluating performance without accounting for the representativeness of the selected labeled data. Furthermore, some studies have used inappropriate or limited metrics. This motivates our investigation into fairer evaluation metrics for AL algorithms. In this study, we review current AL evaluation metrics, highlight underexplored but emerging ones, and outline key research questions. We also aim to highlight previously unexplored evaluation metrics that could provide valuable insights into evaluating active learners from perspectives not addressed by traditional metrics.

Read the paper · More papers on PaperTik