Calibration Degradation in Gradient-Boosted Trees Versus Deep Tabular Neural Networks Under Subpopulation Shift and Class Imbalance: An Empirical Study with Post-Hoc Recalibration
Ahmad Raza · Zenodo (CERN European Organization for Nuclear Research) · 2026
Reliable probability estimates are as important as raw discriminative accuracy in high-stakes tabularclassification settings such as lending, hiring, and public-benefit screening, yet most comparisonsbetween gradient-boosted decision trees (GBDTs) and deep tabular neural networks focus onaccuracy or AUROC and leave calibration largely unexamined outside of the training distribution. Thispaper reports a controlled empirical study of predictive calibration for a LightGBM gradient-boostedclassifier and an embedding-based deep tabular neural network (an entity-embedding multilayerperceptron), benchmarked against a logistic regression reference, on the UCI Adult Census Incomedataset (48,842 records after cleaning). Models were trained only on a United States-residentsubpopulation and evaluated under three conditions: (i) an in-distribution (ID) held-out test split, (ii)a set of controlled class-imbalance variants obtained by subsampling the positive class of the ID testsplit to positive rates between 24% and 2%, and (iii) an out-of-distribution (OOD) subpopulation-shifttest split consisting exclusively of foreign-born individuals, none of whom appear in training. Post-hocrecalibration (Platt scaling, isotonic regression, and, for the neural network, temperature scaling) wasfit exclusively on in-distribution validation data and evaluated for transfer to the shifted domain.At ID, GBDT and the deep tabular network are statistically indistinguishable in calibration (ExpectedCalibration Error, ECE = 0.0115 vs. 0.0099; bootstrap 95% CI of the difference [-0.0058, 0.0074], p =0.84) while GBDT retains a significant AUROC advantage (0.929 vs. 0.911, p < 0.001). Undersubpopulation shift, however, GBDT calibration degrades far less than the neural network's: ECE risesto 0.0206 for GBDT versus 0.0352 for the deep network, a statistically significant gap (95% CI of thedifference [-0.0183, -0.0057], p = 0.002). Controlled imbalance experiments show ECE increasingmonotonically for all three models as the positive rate falls from 24% to 2% (up to a 9-fold increase inECE), with GBDT consistently the most robust. Critically, post-hoc recalibrators fit on in-distributionvalidation data and transferred to the shifted domain did not uniformly help: Platt scaling nearlydoubled ECE for every model under shift (e.g., GBDT: 0.0206 to 0.0402), while isotonic regression andtemperature scaling were comparatively neutral to mildly beneficial. These results indicate that rawdiscriminative performance and in-distribution calibration are each poor proxies for calibrationrobustness under realistic subpopulation shift, and that naively transferring a validation-fitrecalibrator to a shifted population can actively worsen the reliability of predicted probabilities. Thefull experimental pipeline, trained artifacts, and figures are reported for reproducibility.