DSFB-GPU: Clear-Box Pure Deterministic Inference CUDA Acceleration for Replayable Trace-Event Verdicts A Prior-Art Architecture for non-probabilistic, non-stochastic, non-weighted, GPU-Accelerated Residual Signs, Detector Motifs, Bank-Governed Fusion, and Byte-Exact Case Files Without Probabilistic Models
Riaan De Beer · Zenodo (CERN European Organization for Nuclear Research) · 2026
Modern AI acceleration is dominated by probabilistic black-box inference: the workloadthat the hardware accelerates is a learned, weight-bearing model whose internal route fromevidence to conclusion is opaque to operators. Regulated, debug, and industrial environmentsincreasingly require the opposite: high-throughput inference whose reasoning path is itselfdeterministic, inspectable, replayable, and verdict-producing. This paper presents DSFB-GPU a bounded prior-art proof of clear-box deterministic inference accelerated on GPUhardware. The pipeline maps the seven canonical stages of Dsfb debug inference (de Beer,2026m) — residual extraction (de Beer, 2026i), drift/slew sign construction (de Beer, 2026|),detector motif scoring, consensus grid formation, candidate collapse, bank-governed episodeemission, and replayable case-file assembly — onto fixed-point CUDA kernels and a Rust-resident heuristics bank. The GPU accelerates only the evidence-production stages; the bankretains semantic authority. No neural network is used, no probability distribution is consulted,no learned weight is invoked, no stochastic sample is drawn. The result is not a prediction; it isa replayable verdict case file whose hash chain anchors every intermediate artifact back to theinput catalog. We name this inference mode endoduction (§2.1): deterministic adjudicationof internal evidence-field relations into a replayable structural verdict, distinct from induction,deduction, and abduction. We demonstrate byte-exact CPU/GPU equivalence at everystage on a synthetic fixture, extend the same court contract over 20 vendored real-datafixtures spanning 5 source-class families (the S-REAL audit gauntlet, 316 admitted episodes,byte-identical replay across two within-run dispatches per dataset), classify 30 fixtures bysaturation regime (10 saturation-class real-data fixtures up to 117 % of the S-PERF.16.asynthetic throughput anchor; 20 launch-bound fixtures honestly reported), and publish apublic Colab replay surface (COLAB.S-REAL.1) so external evaluators can rebuild theCUDA path from source and re-run the audit against the plan-locked vendored fixtures witha recorded cross-hardware byte-identity verdict.Keywords: DSFB, DSFB-GPU, Drift-Slew Fusion Bootstrap, deterministic inference, clear-box inference, endoduction, densorial inference, tekmeric inference, deterministic evidence court, CUDA acceleration, GPU evidence factory, replayable case files, byte-exact replay, fixed-point inference, Q16.16 arithmetic, hash-chained receipts, Semantic Non-Bypass Axiom, residual densors, detector motifs, bank-governed fusion, deterministic witness families, trace-event inference, observability telemetry, debug telemetry, real-data audit, public reproducibility, Colab replay, Nsight Compute, GPU measurement discipline, prior art, Densor Processing Unit, DPU architecture, deterministic AI, non-probabilistic inference, non-stochastic inference, non-neural inference, reproducible GPU computing, evidence densor, DSFB-GPU-Atlas