Silent Data Corruption--Limited Scaling Kinetics for Large-Scale AI Training
K Takahashi · Zenodo (CERN European Organization for Nuclear Research) · 2025
Silent data corruption (SDC) is becoming a practical limiter for large-scale AI/LLM training: uncorrected GPU memory faults (HBM/ECC edge cases), storage bit rot, network/collective miscompares, and software/hardware divergences that do not reliably crash jobs, yet can silently alter gradients, optimizer state, checkpoints, or metadata. At fleet scale, “rare” integrity events become routine—and conventional protection layers can be insufficient without explicit instrumentation and disciplined recovery. This paper proposes a telemetry-contract framework that turns integrity checks into verifiable obligations. Instead of treating missing checks as “nothing happened,” the system makes missing evidence a contract violation. The approach defines (1) a tamper-evident, canonicalized evidence log (e.g., RFC 8785 canonical JSON, hash chains / optional transparency logs), (2) a conservative strictness predicate that fails closed on missingness, contradictions, or liveness gaps, and (3) an auditable controller that triggers pause/rollback/quarantine when the run cannot justify its own integrity. Two operational outputs are introduced: certified progress PC(t)PC(t)PC(t) (only progress accrued while strictness holds) and a certified useful compute floor that combines certifiability with telemetry overhead (TOR). Under explicit assumptions (log completeness, collision bounds, and empirically declared “test escape” rates), the paper provides a coverage-in-time residual-risk bound and an offline verifier algorithm for third-party audit. The core claim is practical: scaling kinetics should be modeled not only by available compute, but by the fraction of compute that can be certified under realistic integrity budgets.