Toward More Trustworthy QSAR: A Systematic Discussion on Data Set Partitioning
Shangyu Li, Peizhe Sun · Journal of Chemical Information and Modeling · 2026
With the surge in QSAR model development, concerns about evaluation rigor, particularly regarding the influence of data splitting, have grown. Using five data sets of various sizes, we systematically assessed the effects of random splits (RS), similarity-based splits (SS), and random-seed variability on model generalizability under two scenarios: limited data for chemical screening and standard modeling with ample data. Both the choice of data set partitioning method and the selection of random seeds can substantially affect internal test performance, which may not reliably reflect true predictive capability. Although SS can improve internal test performance in many settings, these gains do not necessarily translate into stronger external generalizability. Moreover, under low sampling ratios, SS may perform worse than RS on both internal and external tests. This challenges the implicit assumption that rational splits optimized for internal performance universally improve model performance. Notably, variability across random seeds was high on internal tests in the smallest data set ( R 2: 0.453–0.783), whereas on the fixed external data set R 2 varied less (0.633–0.672), regardless of applicability domain (AD) filtering. This undermined cross-study comparability and underscored the risk of overly optimistic conclusions. Our findings highlighted that test-set construction must be aligned with real-world application scenarios. Researchers should avoid relying on single or cherry-picked random seeds or unsuitable rational partitioning. Transparent, application-aligned partitioning protocols and AD methods should be employed to emphasize true external generalizability over potentially inflated internal metrics.