Clash of titans on imbalanced data: TabNet vs XGBoost
Róbert Kanász, Peter Drotár, Peter Gnip, Martin Zoričák · 2024
In machine learning, particularly with tabular data, ensemble methods and neural networks stand as the preeminent approaches for predictive modeling. Among these, XGBoost and TabNet have demonstrated remarkable efficacy and interpretability. However, one critical challenge in these methodologies is their performance on imbalanced datasets, a common yet intricate issue in many real-world applications. This research paper proposes novel modifications to TabNet, tailored to enhance its performance on imbalanced tabular datasets. Our methodology introduces multiple loss functions for the TabNet architecture. These modifications improve the models’ sensitivity to minority classes and enhance overall predictive accuracy on imbalanced data. We conducted a comprehensive performance comparison using various synthetic and real-world datasets characterized by significant class imbalances. TabNet, combined with IBLoss, achieved a GM score on real-world data up to 92%. On synthetic data, the highest GM score was up to 82% using TabNet in combination with BVSLoss. These results demonstrate that TabNet is robust to imbalanced datasets and can learn well even on imbalanced data. The performance is further boosted by incorporating a loss function built for imbalanced data, such as BVSLoss or IBLoss. On the other hand, XGBoost fails to converge if not adapted to imbalanced data with sampling or cost-sensitive learning, resulting in less accurate prediction performance.