A Data-Driven Approach for Balancing Overfitting and Underfitting in Decision Tree Models

Mykola Zlobin, Volodymyr Bazylevych · Central Ukrainian Scientific Bulletin Technical Sciences · 2025

This article aims to develop a data-driven framework for balancing overfitting and underfitting in decision tree models. Overfitting occurs when a model captures noise, reducing generalization, while underfitting leads to poor predictive accuracy. The study systematically tunes the max_leaf_nodes parameter and evaluates model performance using Mean Absolute Error (MAE). The objective is finding the most optimal balance that ensures model accuracy while preventing excessive complexity. A Decision Tree Regressor has been trained on the Ames Housing dataset, which includes 79 explanatory variables related to home prices. The dataset has been splitted into training and validation sets. The model has been evaluated by iterating over different max_leaf_nodes values, ranging from 2 to 5000, and computing the MAE for each configuration. The results show that increasing max_leaf_nodes initially improves accuracy, but beyond 400 nodes, MAE stabilizes around 242,906, indicating that further complexity does not improve performance. The paper highlights that models with too few leaf nodes underfit the data, while models with too many leaf nodes overfit, capturing spurious patterns. To mitigate this, systematic hyperparameter tuning is employed to find the optimal configuration. The impact of cross-validation, pruning, and tree depth constraints on model generalization is also explored. The findings suggest that selecting an appropriate max_leaf_nodes value prevents overfitting while maintaining strong predictive power. Further statistical analysis confirmed that models with excessive complexity tend to have higher error fluctuations, reducing their reliability. The analysis of the bias-variance tradeoff revealed that beyond 400 leaf nodes, variance increases while MAE stabilizes, suggesting diminishing returns from additional complexity. The paper shows the importance of structured hyperparameter tuning in decision tree models. The optimal max_leaf_nodes value is found at 400. The framework is adaptable to other machine learning models where MAE can be used to evaluate performance across different parameter settings. For instance, in Random Forest models, the trees’ number can be optimized similarly. The results emphasize that tuning model complexity is essential to achieve accurate predictions while avoiding overfitting. Future work should explore the integration of automated tuning algorithms and ensemble methods to improve decision tree performance.

Read the paper · More papers on PaperTik