From Recursive Scaffolding to Admissibility-First Construction: Mechanism, Stability, and Failure-Mode Decomposition on OOLONG-Pairs

Shawn Kevin Jason · Zenodo (CERN European Organization for Nuclear Research) · 2026

Empirical companion to the RLM paper, testing its theoretical predictions on a Zhang-style OOLONG-Pairs setting. Decomposes recursive scaffolding into its mechanism components — filter-mode GAF, construction-mode GAF, and the admissibility-relevant summary state — and shows empirically that the essential mechanism is not recursion itself but the construction or approximation of an admissibility-relevant state. Abstract Recursive Language Models (RLMs) have recently shown strong empirical gains on long-context reasoning and aggregation benchmarks, especially OOLONG and OOLONG-Pairs. In prior theoretical work, RLMs were analyzed through the admissibility-dynamics framework as a composite mitigation: they extend effective certification depth by offloading context into an external REPL environment, and they partially escape bounded local generation by constructing intermediate summaries through recursive sub-calls. That analysis predicts that RLMs succeed not because recursion is intrinsically sufficient, but because, in favorable cases, recursive scaffolding constructs an admissibility-relevant summary state. This paper reports an empirical decomposition of that prediction on a Zhang-style OOLONG-Pairs setting constructed from the validated trec_coarse OOLONG-synth split. In the canonical 20-query seed-7700 run at approximately 32K context tokens, direct GPT-5 collapsed under the global pair-construction burden, achieving micro-F1 of approximately 0.0019. RLM(GPT-5) recovered the global pair structure with micro-F1 = 0.9064 and macro-F1 = 0.8400. Model-based GAF filtering over RLM output increased precision and improved micro-F1 to 0.9107, although with a nontrivial recall reduction. Model-based admissibility-first GAF construction without RLM achieved the highest non-oracle micro-F1 in that run, 0.9212, but lower macro-F1 than RLM, 0.8196. This result should therefore be read not as categorical dominance over RLM, but as a different precision–recall operating point: higher recall, lower precision, and comparable aggregate F1 without recursive scaffolding. A representative five-query robustness subset (q8, q9, q11, q15, q20) was then evaluated across seeds 7701, 7702, and 7703 to probe seed sensitivity, query-level variance, candidate-generation brittleness, and model-parity concerns. Across the three-seed subset, Z1G-S-mini achieved the best practical micro-F1, 0.9169, while RLM and Z2G-M achieved 0.7304 and 0.7335 respectively. The aggregate masks an important bimodal behavior. On seeds 7701 and 7703, RLM recovered high performance (micro-F1 = 0.9048 and 0.9111). On seed 7702, however, RLM suffered a catastrophic recall collapse (micro-F1 = 0.1379), driven primarily by q8 under-generation; model-based GAF filtering inherited that collapse because it can only reject existing candidates. Construction-mode GAF did not depend on the RLM candidate set and remained high across all three seeds: Z1G-S-mini ranged from 0.9035 to 0.9244 micro-F1. The robustness result therefore separates filter-mode GAF from construction-mode GAF: filtering is valuable when a rich candidate set exists, but construction is robust to candidate-generation failure. GPT-5 construction remains a model-parity and tail-query diagnostic rather than a complete three-seed lane: across seeds 7701 and 7702 it achieved micro-F1 = 0.9217 and the strongest macro-F1, 0.7026. A direct output-structure comparison on q15 seed 7702 provides the strongest mechanism-level evidence. RLM and Z1G-S-Big produced identical aggregate counts but not identical pair sets; instead, they shared a large role-product core and differed by a symmetric single-endpoint false-positive substitution. This indicates that the recursive scaffold and direct admissibility classifier converged to the same admissibility-summary form, with disagreement isolated to endpoint-level classification noise. The updated analysis also distinguishes endpoint-classifier noise, false-positive amplification, residual pair-level relation structure, verifier miscalibration, candidate-generation collapse, benchmark/seed validity, and cost/routing as separate bottlenecks with different mitigations. The central conclusion is mechanism-level. On OOLONG-Pairs, where pair validity is often factorizable through endpoint admissibility but can expose residual pair-level constraints, recursive scaffolding is not the essential mechanism. The essential mechanism is the construction or approximation of an admissibility-relevant state. RLM is one route to that state; GAF-style admissibility-first construction states the mechanism directly. The full experimental sequence reported here remained at small-lab scale, under USD $100 in API charges in this implementation, making the bridge experiment reproducible without institutional compute. The results remain limited by adapted benchmark construction, partial multi-seed coverage, synthetic grouping, single-run conditions per seed/query, and the favorable factorizable and partially factorizable structure of OOLONG-Pairs. They should be interpreted as directionally strong empirical bridge evidence rather than a statistically powered benchmark claim. Companion Lean 4 formalization: https://doi.org/10.5281/zenodo.20062396 GitHub repository: https://github.com/shawnjason/OOLONG-Pairs Related papers in the program: PIT (foundational projection-theoretic result): https://doi.org/10.5281/zenodo.19633241NEO (forward-case impossibility theorem): https://doi.org/10.5281/zenodo.19688367IA (admissibility-dynamics framework): https://doi.org/10.5281/zenodo.19688628HAL (language-model specialization): https://doi.org/10.5281/zenodo.19715059RLM (theoretical predecessor whose predictions this paper tests empirically): https://doi.org/10.5281/zenodo.19753549SUD (Sudoku-Microscope empirical validation): https://doi.org/10.5281/zenodo.20277939HAM (Hamiltonian-Microscope cross-provider pilot): https://doi.org/10.5281/zenodo.20278073

Read the paper · More papers on PaperTik