Axiom-World: A Pre-Registered Study of Two-Stage Post-Training for Rule-Grounded Planning in a Verifiable Toy World

Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Bi, Xiao, Zhang, Haowei, Mingchuan Zhang, Y. K. Li, Yingjun Wu, Daya Guo · arXiv (Cornell University) · 2024

Abstract We study whether a two-stage post-training recipe — general reasoning SFT (Phase 1) followed by task-specific tuning (Phase 2) — outperforms direct task tuning for rule-grounded planning, using a fully verifiable synthetic environment (PlayWorld) with frozen, fingerprint-pinned evaluation suites spanning in-distribution, three OOD axes, and adversarial traps. All arms, comparisons, champion-selection rules, and stop rules were pre-registered before the first run; every deviation is logged as a numbered amendment. Findings. (1) The two-stage champion (B4v2: general-reasoning SFT → PlayWorld SFT) beats the direct-tuning control (A2v2) on all five suites across all three seeds (Δpass +10 to +22 pp; paired permutation p ≤ 0.0004 per seed per suite; sign-consistent 15/15). (2) Preference optimization (DPO) on top of either route is approximately null, while the two-stage advantage fully survives it (B5 ≫ A2v2, 10/10 suite-metrics, p ≤ 0.0002). (3) Verifier-rewarded online RL (GRPO) regressed the champion under two reward designs; diagnostics attribute the failure to advantage starvation (50–70 % zero-variance reward groups) and entropy collapse (0.15 → 0.008) producing required-component specialization on a narrow training scenario pool — reward gating recovered only the adversarial suite (net +15/300 episode flips), confirming the mechanism. (4) Off-the-shelf instruct baselines fail primarily at the format gate (0-shot: ~96 % malformed JSON), while few-shot prompting recovers adversarial trap-avoidance (.907) but not construction (ID .227). We release all run artifacts with sha-pinned lineage (base-model revision, dataset fingerprints, parent-adapter hashes).

Read the paper · More papers on PaperTik