Design-based conformal prediction
Wieczorek, Jerzy · arXiv (Cornell University) · 2023
Abstract Split conformal prediction converts any model into a distribution-free predictor whosecoverage, conditional on the calibration set, is a Beta-distributed random variable ratherthan a fixed number [1,2]. That characterisation assumes i.i.d. calibration data. We derive itsgeneralisation when calibration units instead arrive in correlated families — beam branchessharing a prefix, resampled particles sharing an ancestor, retrieved calibration sets rebuiltper query. For b i.i.d. clusters of m within-cluster-exchangeable scores, coverage is asymptoticallynormal with variance p(1-p)/n_eff at an effective sample size n_eff = n/[1+(m-1) rho_I(p)] [3],where rho_I(p) is the intra-cluster correlation of the exceedance indicators [4,5] at thecoverage level p — not of the scores. The i.i.d. Beta law with n replaced by n_eff matchesthose two moments and is the form we use throughout, as an approximation rather than acorollary: a central limit theorem fixes two moments, not a distribution. The law depends onthe score distribution only through the copula diagonal delta(p), the probability that twosame-cluster scores both fall below the p-quantile [6,7]; the density cancels, so the result isdistribution-free given the copula [8,9,10]. It recovers Beta(k, n+1-k) at zero correlation[1,2] and the b-cluster law at perfect correlation [11], closing an interpolation named as anopen problem in 2023 [12] and left underived by two 2026 preprints [11,13]. Unequal familysizes are handled by replacing m with the size-biased mean m_tilde = sum_j m_j^2 / sum_j m_j,which for realistic beam-search profiles is roughly twice the average family size — so theintuitive correction is less than half the required one. Four consequences. First, the design effect that the recent literature writes down — Kish'sformula applied to the score correlation [13] — is wrong, with 5.7x the error of the correctlaw across our range, and it is not conservative by design: we exhibit a valid copula on whichthe score correlation is exactly zero, so the correction vanishes, while the true design effectis 1.44. Second, the correct design effect is level-dependent and weakest in the tails (1.97 at90% coverage, 1.54 at 99%, same rho), so the harm from shared ancestry is not a single number asystem can report once. Third, a second-order term biases coverage below nominal at rate(m-1){p(1-p) rho_I'(p) - (2p-1) rho_I(p)}/2n — and level-dependence is exactly what fixes itssign, so the property that makes clustering cheaper in variance is the property that makes itunsafe in mean. Fourth, because the dispersion penalty is first-order while the mean bias isonly O(1/b), the damage is nearly invisible in the statistic these papers report. We audit eight LLM inference-time systems for how they obtain exchangeability, find fivestrategies of which two are defensible, and measure the effect on a released calibration set —by resampling it, not by evaluating our own formula on it: a process-reward-model system whose25,028 calibration points [14] deliver coverage dispersion 4.4x wider than the exchangeabilityit assumes would imply, once the score's ties are broken; on the raw nine-atom score theinflation is nearly invisible, because the score violates the theorem's own continuityassumption. The plug-in design effect over-predicts the measured inflation by a fifth, and thediscrepancy is itself informative: the size-biased correction over-corrects when family sizesare informative. In the same artifact we find a violation larger than the one we characterise. Family size isstrongly coupled to family score (Spearman -0.42, p ~ 1e-23): harder questions produce longerreasoning traces, hence more prefixes and lower success probabilities. That coupling breaks thecommon-marginal assumption at first order, and how much it costs turns on a reading thesystem's authors never state — under a per-question test marginal it moves coverage by 4.99percentage points, 69 times the mean effect the design effect produces on the same data, andunder a per-prefix marginal it is exactly zero. Neither reading is written down, so neither canbe ruled out. We report this against our own interest, because a design-effect correctionapplied to data with informative cluster sizes would be a precise answer to the wrong question. What changed in v5 Eight corrections. Four of them change something a reader of an earlier version may be actingon. 1. The headline is measured, not derived. Coverage dispersion on the released PRM calibration set is 4.4x wider than exchangeability implies, and its 25,028 points carry roughly 1,300 points of information — not 5.6x and roughly 800. The old figures were the square root of the plug-in design effect, the paper's own formula evaluated at its own estimate, inside a section that presents itself as a measurement. v5 cluster-bootstraps the released questions through split conformal against an i.i.d. control. On the raw score the inflation is 1.09x, because the score takes nine values with 67% of mass at zero and so violates the continuity assumption; random tie-breaking is the standard repair and reveals the 4.4x. 2. Theorem 1 loses its Beta clause. Through v4 the theorem box said "equivalently, coverage follows the i.i.d. Beta law at the effective sample size". A central limit theorem fixes two moments, not a distribution; the Beta substitution matches mean and variance and nothing further, and no rate has been proved for it. Section 3.2 now displays the normal limit and gives the Beta form underneath as a labelled approximation. Nothing that was proved has been withdrawn — what changed is which half is claimed as proved. 3. A Section 6.1 caveat is withdrawn as wrong, and it corrects us against our own interest. Earlier versions said the release carries no trajectory index, so the reported design effect was a lower bound. The index is recoverable: the system takes every prefix of each trajectory, so one trajectory's prefixes nest under string-prefixing, giving 3,961 maximal chains against 4,000 nominal with 97.9% of rows on exactly one chain. At that finer level the exceedance correlation is indeed higher, 0.688 against 0.495 — but the design effect is smaller, 7.06 against 30.8, because the size-biased mean family size falls from 61.3 to 9.8 and the design effect is dominated by family size rather than correlation. Since the system samples questions, 30.8 is the complete number, not a floor beneath a larger one. 4. Two scope conditions and two attributions. Proposition 2's size-biased-mean substitution over-corrects when cluster sizes are informative and now carries that condition; Corollary 3's sign claim is stated with its hypotheses rather than as a consequence of the coverage level; and Section 3.1 no longer claims the exceedance correlation is below the score correlation "always". Liu, "Target-indexed hierarchical conformal prediction under informative cluster size", anticipates in general form the test-marginal distinction of Sections 3.5 and 6.1 and is now cited there. Vejling, Biscio, Mazoyer, Popovski and Pandey state calibration-conditional coverage as a Beta law at Kish's weighting effective sample size, so any suggestion that no one in the field computes a coverage-variance effective sample size is withdrawn; the accurate statement, now in Section 2, is that the conformal literature computes it for weights while assuming away dependence within the calibration set. Also in this version: Section 3.5 credits Einbinder et al. and Sesia et al. for thesame-rate-opposite-sign phenomenon and the governing sign identity; Figure 4 and the herographic carry the measured numbers. The verification archive (32 files) gains prm_dispersion.py,the measurement behind item 1, and trajectory_index.py, the reconstruction behind item 3, eachwith its recorded output. Cite the concept DOI, 10.5281/zenodo.21595640, which always resolves to the newest version. References The works the abstract above depends on directly, numbered in order of first appearance. Thepaper's full bibliography — 47 entries, each annotated with what was read and at what depth —is in the PDF and in the verification archive. [1] Vovk, V. (2012). Conditional validity of inductive conformal predictors. ACML; Machine Learning (2013). https://proceedings.mlr.press/v25/vovk12.html — coverage conditional on the calibration set is Beta-distributed, not a fixed number. This paper's zero-correlation endpoint. [2] Bian, M. & Barber, R. F. (2022). Training-conditional coverage for distribution-free predictive inference. Electronic Journal of Statistics. arXiv:2205.03647 — joint owner with [1] of that endpoint. [3] Kish, L. (1965). Survey Sampling. Wiley, New York. — deff = 1 + (m-1) rho for means and proportions. The form Theorem 1 takes; what is new here is which correlation belongs in it. [4] Kraemer, H. C. (1979). Ramifications of a population model for kappa as a coefficient of reliability. Psychometrika 44(4):461-472. https://doi.org/10.1007/BF02296208 — the binary intra-class correlation and its exact relation to the continuous one under a normal latent model. The exceedance correlation, and its dependence on the cut point, are his; in print since 1979. [5] Donner, A. & Eliasziw, M. (1994). Statistical implications of the choice between a dichotomous or continuous trait in studies of interobserver agreement. Biometrics 50(2):550-555. https://doi.org/10.2307/2533400 — restates [4] as their Eq. (3), with the bound and the observation that the divergence grows as prevalence approaches 0 or 1. [6] Venter, G. G. (2002). Tails of copulas. Proceedings of the Casualty Actuarial Society LXXXIX:68-113. — the tail concentration function built from the copula diagonal. Section 5.1 shows rho_I(p) is an affine reparameterisation of it. [7] Durante, F., Fernández-Sánchez, J. &