Quantum Optimization and Reinforcement Learning for Real-Time Plasma Control: QUBO/QAOA Actuator Scheduling and Quantum RL for Shape Control, without Claimed Advantage
Priyanca Ford, G L Kulcinski · Zenodo (CERN European Organization for Nuclear Research) · 2026
Real-time control of a compact spherical-tokamak breeder poses, at several points, a combinatorial-optimization problem inside a control cycle: which actuators to schedule, how to lay out a discrete blanket, and — the control-relevant case — how to sequence poloidal- and toroidal-field coil currents under hard constraints while tracking a shape and vertical-position target. Because these problems are naturally binary and quadratic, they are the canonical near-term target for quantum optimization, and it is a fair, testable question whether a quantum processor solves them faster or better than a classical one. This paper answers that question with a formal treatment and a pre-registered result. We derive compact quadratic-unconstrained-binary-optimization (QUBO) encodings for all three problems from first principles — pulse cadence as a number-partitioning Ising model, blanket layout as a penalty-embedded assignment, and coil-current/actuator scheduling as a fixed-point-encoded quadratic tracking objective built on the free-boundary equilibrium shape-response Jacobian — and map each to two fault-tolerant primitives, a depth-three quantum approximate optimization ansatz (QAOA, p=3) and a D\"urr–Hyer Grover minimum-finder. At the frozen breeder design point (I_p=9.66 MA, Q=3.076, P_fus=85.04 MW, negative triangularity δ=-0.30, B_0=8 T, κ=2.0, =1.20 m, aspect ratio A=2.5, centrepost peak field 16.84 T) the design-scale instances need 23–61 logical qubits and 5×10^3–5.5×10^7 T-gates, verified against the Fowler surface-code and D\"urr–Hyer formulas across 50 resource rows with zero discrepancies. On the question of whether quantum hardware helps, we report an explicit negative: classical simulated annealing matches or beats QAOA on all three breeder QUBOs out to n=120 variables; the exact-to-heuristic crossover sits at n=12–16 (exact optima in 0.03–292 s of wall time); the QAOA approximation ratio plateaus at 0.90–0.95 and the circuits become classically unverifiable past n≈24; and a fleet-scale (n=120) Grover formulation would need 1.8×10^18 oracle calls and 1.2×10^23 T-gates — a resource boundary, not a route. We therefore place variational quantum reinforcement learning for shape and vertical-position control on a dated, benchmark-gated schedule: a quantum policy is admitted only if it meets or beats the classical controller at equal wall-clock with a control-barrier-certified safety filter held inviolate. All quantum and classical runs were carried out on the NVIDIA GPU HPC campaign. The deliverable is the honest, resource-grounded map of where quantum computing does — and does not — pay for real-time fusion control. Key results (frozen anchors): N_T = 1.2 ×10^23. Methods & codes: FreeGSNKE, FreeGS, Sauter, QAOA, VQE, ADAPT-VQE. Live verification: 4 gate validator(s) with live-recompute cards (2 reproduced, 1 revised, 1 estimated). See the Verification section and data/verification.csv. Related: Paper page · De-risking register · 3D model · Learn more about Kronos Part of the 2026 Kronos publication series; independently re-run and stamped in the Kronos de-risking register (DOI 10.5281/zenodo.22645689). All numerical values are frozen design-point anchors; see the register. Public research artifact. No proprietary, financial, or supply-chain information is included.