A Model-Agnostic, Performance-Pushforward Theory of Scaling Laws
Takahashi, K · Zenodo (CERN European Organization for Nuclear Research) · 2025
This paper proposes a unified, model-agnostic mathematical theory for LLM and foundation-model scaling laws. The central idea is to view performance as the pushforward of an EVIλ_\lambdaλ gradient flow under a Lipschitz evaluation map, while scale is measured by the geometry of preimages of performance targets. Working on the image quotient metric (no curvature loss) and, in parallel, on external image metrics (with the sharp degradation λ↦λ/L2\lambda\mapsto \lambda/L^2λ↦λ/L2), we derive tight time-to-accuracy bounds and compute-optimal allocation rules that are independent of the model family. Main contributions (all with full proofs): Image EVI on the quotient (fibre-invariance): Existence of a projected semigroup VtV_tVt with EVIλ_\lambdaλ on the image pseudometric dOd_OdO and e−λte^{-\lambda t}e−λt contraction. External image EVI (general LLL-Lipschitz observers): A complete proof of EVI with curvature degradation λ/L2\lambda/L^2λ/L2 using a fibrewise inf-projection of the energy. Energy–speed limit (lower bound): A quadratic length–energy inequality yields τ(ε)≥σx0(ε)2/(2ΔE)\tau(\varepsilon)\ge \sigma_{x_0}(\varepsilon)^2/(2\Delta\mathcal E)τ(ε)≥σx0(ε)2/(2ΔE), clarifying time-to-target lower bounds even when λ=0\lambda=0λ=0. Preimage geometry ⇒\Rightarrow⇒ exponents: If preimages are totally bounded with box-counting dimension dpre(ε)d_{\rm pre}(\varepsilon)dpre(ε), then quantization error scales as O(N−1/dpre)O(N^{-1/d_{\rm pre}})O(N−1/dpre) and Rademacher complexity as O~ (dpre/D)\tilde O\!\big(\sqrt{d_{\rm pre}/D}\big)O~(dpre/D). Proofs include measurability and chaining details. Compute-optimal allocation (with discretization): For a FLOPs ansatz F≍κNγNDγDLcγc τ/hF\asymp \kappa N^{\gamma_N}D^{\gamma_D}L_c^{\gamma_c}\,\tau/hF≍κNγNDγDLcγcτ/h and balanced errors N−α ≍ D−β ≍ e−cτ ≍ h2+β′N^{-\alpha}\!\asymp\!D^{-\beta}\!\asymp\!e^{-c\tau}\!\asymp\!h^{2+\beta'}N−α≍D−β≍e−cτ≍h2+β′, we obtain closed-form exponents that explicitly include the discretization term α/(2+β′)\alpha/(2+\beta')α/(2+β′) (recovering classical formulas in the fixed-hhh regime). FBH–ET dictionary: Identification αf=1/(4f′′(1))\alpha_f=1/(4f''(1))αf=1/(4f′′(1)) for the dynamic–static duality (Bures/HK/BKM/KL at quadratic order), with scope conditions stated. Strang splitting with metric BCH: Global O(h2)O(h^2)O(h2) accuracy under a resolvent-level metric BCH bound; includes an “audit” recipe for safe operator splitting. Multimodal extension: Vector-valued evaluations with an ℓp\ell_pℓp product metric; pulled-back product pseudometric d⊕d_\oplusd⊕ controls covering numbers and yields a dimension upper bound dpre∏≤∑jdjd_{\rm pre}^{\prod}\le \sum_j d_jdpre∏≤∑jdj. Aggregators Ψ\PsiΨ are Lipschitz-calibrated (temperature/margin) to prevent spurious emergent effects. Practical bridge (Appendix A, theory-first):Constructive, experiment-light procedures to estimate λeff\lambda_{\rm eff}λeff (sublevel contraction), dpre(ε)d_{\rm pre}(\varepsilon)dpre(ε) (doubling/KNN slope), and Lipschitz constants for evaluations and aggregators, enabling reproducible, model-family–agnostic scaling design. What this delivers for practitioners A unified scaling law blueprint: swap architectures without changing the calculus—only the evaluation’s Lipschitz constants and preimage geometry enter. Actionable bounds on training time and accuracy that separate optimization, approximation, statistics, and discretization. Compute planning that natively handles context length LcL_cLc and non-unit compute exponents (γN,γD,γc)(\gamma_N,\gamma_D,\gamma_c)(γN,γD,γc). Document quality: OCR/crawler-friendly PDF (selectable text), 1.3 line spacing, hyperlinked references, full proofs, and notational glossary. Keywords: LLM scaling laws; gradient flows in metric spaces; Ambrosio–Gigli–Savaré (EVI); pushforward performance; image pseudometric; quotient metric; fibre/fiber invariance; energy–speed limit; preimage geometry; covering numbers; Minkowski/box-counting dimension; Rademacher complexity; compute–data–model allocation; discretization exponent; Strang splitting; metric BCH; Bures/HK/WFR; BKM; KL; multimodal alignment; Lipschitz calibration. Suggested communities/tags: Machine Learning; Optimization; Information Theory; Numerical Analysis; Theoretical Computer Science; Statistics & Learning Theory.