Re: Don’t Let Your Analysis Go to Seed: On the Impact of Random Seed on Machine Learning-based Causal Inference
Nicholas Williams, Anton Hung, Kara E. Rudolph · Epidemiology · 2025
To the Editor: Schader et al.1 highlight an important issue that estimators that use machine learning in model fitting can suffer from random seed dependence. However, it is possible that the message may be unnecessarily alarmist, as Schader et al.1 limit their simulation to only using two-folds for cross-fitting. To illustrate our point, we performed a brief simulation study using the same high-dimensional data-generating mechanism from Schader et al.1 We generated j:1,...,200 datasets of size n = 100. For each dataset j, we estimated the average treatment effect 100 times (using 100 different initial random seeds) using an ensemble of main-terms generalized linear models, multivariate adaptive regression splines, and random forests. We performed this simulation under two different cross-fitting parameterizations: 2-folds and 40-folds. The code for our simulation is available on GitHub (https://github.com/CI-NYC/random-seed-effects). Our results are summarized in Figures 1–3. We find that simply increasing the number of cross-fitting folds to an appropriate level for the dimensions of the dataset resolves the issue of estimate stability. The median within-dataset estimate variance (scaled by 104) with two-folds was 0.398 compared with 0.00813 with 40-folds (Figures 2 and 3); the variance of the empirical distribution of the within-dataset estimate variance was 98% smaller with 40-folds compared with two-folds.FIGURE 1.: Replication of Figure 1 in Schader et al.1 with the cross-fit folds increased to 40.FIGURE 2.: Boxplots of the within-dataset ATE point-estimates for each generated dataset. ATE indicates average treatment effect.FIGURE 3.: Histograms of the empirical variance of the within-dataset ATE point-estimates. ATE indicates average treatment effect.Schader et al.1 advocate for the use of averaging estimates over multiple runs of an estimator as a solution to the issue of estimate stability due to random seed dependence. While doing so is a valid approach, it is also likely unnecessary. Cross-fitting is already recommended, which Schader et al.1 summarize, as it mitigates the risk of over-fitting2; if cross-fitting is not used appropriately, then one needs to invoke a Donsker class assumption, which basically states that the estimator function is not too complex.3 As such, choosing an appropriate number of folds for cross-fitting is likely a simpler solution that requires no extra work on the part of the analyst. Guidance on choosing the number of cross-fitting folds that balances estimator stability with computational cost should be an area of future research. ACKNOWLEDGMENTS We thank David Benkeser and Iván Díaz for their thoughtful feedback.