Dynamic‐Balancing AutoML for Imbalanced Tabular Data With Adaptive Resampling and Complexity‐Aware Analysis
Marcelo Vinícius Cysneiros Aragão, Tiago de M. Pereira, Mateus de Freitas Carvalho, Felipe A. P. de Figueiredo, Samuel Baraldi Mafra · International Journal of Intelligent Systems · 2025
Handling class imbalance is a fundamental challenge in supervised learning, particularly in real‐world scenarios where minority classes are critical yet underrepresented. This paper presents a novel dynamic‐balancing pipeline that enhances automated machine learning (AutoML) performance on imbalanced tabular datasets. The proposed approach integrates both traditional and generative resampling techniques with adaptive, class‐specific thresholds, enabling automated and dataset‐sensitive balancing strategies. To assess its generalizability, the pipeline is applied uniformly across binary, multiclass, and multilabel classification tasks. Each configuration is evaluated within an AutoML framework using performance and efficiency metrics, with outcomes validated through statistical testing and effect size analysis. The study also incorporates dataset complexity measures—including feature‐label dependency and class overlap—to investigate how structural characteristics affect balancing efficacy. By combining principled resampling, exhaustive grid search, and rigorous evaluation, the pipeline enables more robust and efficient AutoML workflows. This work contributes a flexible and reproducible framework for addressing class imbalance, particularly in multilabel contexts, and establishes a foundation for scalable, complexity‐aware resampling in automated model development.