Active Learning for Imbalanced Classification: Empirical Insights, Iteration Scheduling, and LLM-Augmented Validation
Lucas H. Benevides E Braga · IEEE Access · 2025
In many real-world machine learning applications, obtaining labeled data is costly and time-consuming, particularly in domains such as medical diagnostics, fraud detection, and customer lead qualification. Active Learning (AL) mitigates this challenge by selectively querying the most informative samples for annotation, thereby improving label efficiency. This work presents a reproducible, statistically rigorous evaluation of three core AL strategies, uncertainty sampling, diversity-based sampling, and query-by-committee (QBC), on two public datasets: the Bank Marketing dataset and the European Credit Card Fraud dataset. Four model–preprocessing configurations were tested over 10 independent runs with 11 AL iterations each, and compared to passive learning baselines using paired t-tests, Wilcoxon signed-rank tests, and effect size analysis. The best configuration, a regularized Logistic Regression with standardized features, achieved statistically significant relative gains of +6.57% in F1-score and +7.34% in accuracy. On the Credit Card Fraud dataset, the best LightGBM configuration achieved an F1-score of 0.83 compared with 0.39 under passive learning, more than doubling relative performance. Analysis of iteration scheduling revealed that random initialization followed by early uncertainty sampling, mid-course diversity sampling, and a final QBC round produced the most consistent improvements. In addition, we demonstrate a complementary large language model (LLM)-based prompt validation workflow on the NYC Restaurant Inspection dataset, showing that Perplexity Sonar achieved the highest F1-score (97.6%) while OpenAI GPT-4o and Gemini 1.5 Pro offered different trade-offs in cost and speed.