Building Trustworthy AI for mRNA Medicine Design: RNA Structure Data Is Redundant, and AI-Optimized Sequences Require Independent Validation

Danielmartin Arogyasami · Zenodo (CERN European Organization for Nuclear Research) · 2026

Messenger RNA therapeutics — including the COVID-19 vaccines and the individualized cancer vaccines now in registrational trials — depend on the 5′ untranslated region, a short leader sequence that governs how much protein each mRNA molecule produces. Because manufacturing cost scales with dose, designing better leaders is both a biological and an economic problem, and machine-learning models are increasingly used to predict which sequences will work. This paper asks a question that has been assumed rather than tested: does supplying such a model with precomputed RNA folding information make it more accurate? I answer it with an ablation designed to be exact. I train a Transformer encoder on the Sample 2019 massively parallel reporter assay in which self-attention is conditioned on a base-pairing probability matrix (BPPM) through a learned per-block scalar bias — an architecture that is a strict generalization of its own sequence-only ablation, initialized so that structure conditioning is learned rather than enforced, and differing from it by twelve parameters. The model uses the structural signal, and gains almost nothing from it. Adding a real BPPM improves in-distribution R² from 0.9484 ± 0.0007 to 0.9494 ± 0.0010, an effect of the same order as seed-to-seed variation and roughly one twentieth of the gain from replacing the convolutional baseline with a Transformer (+0.0211). Out of distribution the effect is incoherent: across three cell backgrounds and three seeds the sign of the change is consistently negative on one split, consistently positive on another, and inconsistent on the third. That the pathway is nonetheless live is shown by two perturbation controls. Retraining on shuffled or permuted BPPMs — corruptions that preserve the matrix's statistics but destroy the sequence-to-structure correspondence — costs 0.0096 [0.0086, 0.0107] and 0.0031 [0.0021, 0.0041] R² respectively, placing both below the no-BPPM baseline, and the learned gains collapse from mean |gain| 0.118 on real matrices to 0.015 and 0.010 on corrupted ones. The model distinguishes real structure from fake, attends to the former, switches off the latter — and is no more accurate for it. A wrong prior costs roughly ten times more than a right one gains, which means the practical risk of adding a structural prior exceeds its benefit whenever the folding model may be wrong. I attribute this to redundancy rather than irrelevance: a BPPM is a deterministic function of the sequence and can carry no information the sequence lacks. Consistent with this, the same predictor independently recovers translation biology it was never given, including upstream-AUG Kozak sensitivity, its opposite-signed counterpart at the main AUG, and reading-frame-dependent uORF repression that deepens from 1.023× to 0.698× when an in-frame stop converts an N-terminal extension into a genuine upstream ORF. Separately, I characterize a failure mode with direct consequences for AI-guided therapeutic design: cross-predictor agreement is at or below chance in the top-100 band in both directions between two independently trained oracles, and unconstrained gradient design satisfies elementary biological-validity filters less often than random sequence (0.8% versus 39.8%). Independent cross-model validation and hard biological filters are therefore load-bearing components of an mRNA design pipeline, not formalities. Code, weights and benchmarks are released under Apache-2.0. Keywords: 5′ UTR, ribosome loading, RNA secondary structure, attention bias, redundant priors, inverse design, oracle hacking, mRNA therapeutics.

Read the paper · More papers on PaperTik