Overvalidation in Machine Learning: Empirical Evidence and Theoretical Insights

Fabrizio Mori, Antonio Emanuele Ciná, Davide Anguita, Fabio Roli, Luca Oneto · IEEE Access · 2026

Over the past decades, advances in Machine Learning have greatly increased both the variety and the complexity of available algorithms. As a result, solving specific tasks often requires making numerous decisions, such as selecting the most suitable algorithm, designing an appropriate architecture, and tuning the corresponding hyperparameters. Although numerous methods have been proposed to assist this decision-making process, the final selection is typically based on performance over a holdout set. While this practical approach is widely adopted and generally effective, it may give rise to the problem of overvalidation when the number of possible choices becomes large. Overvalidation refers to the bias in holdout performance estimates, which may lead to either a selection bias, that is, the choice of a suboptimal model, or a generalization bias, that is, an overly optimistic estimate of the selected model’s performance. This issue can be mitigated by increasing data quantity and quality, by improving resampling strategies, or by carefully reducing the number of choices. However, both theoretical and empirical evidence show that it tends to reemerge as the number of choices grows, or due to data contamination and poorly designed ML pipelines. This problem has been known for many years, and several researchers have demonstrated these biases in both specific and general scenarios, while also investigating some of the underlying theoretical aspects. Nevertheless, the problem is still only partially analyzed and not yet fully understood. For this reason, in this work we conduct a series of empirical evaluations to demonstrate the presence of overvalidation across different datasets and state-of-the-art architectures. Furthermore, we leverage statistical learning theory to shed light on the empirical evidence, showing how it can help in both detecting and mitigating this phenomenon.

Read the paper · More papers on PaperTik