Axiom: A Householder-Parameterized Pure Unitary RNN for Long-Range Sequence Modeling

Chaudhary, Sanyam · Zenodo (CERN European Organization for Nuclear Research) · 2026

We present **Axiom**, a recurrent neural network whose hidden-to-hidden transition is parameterized as a product of Householder reflections, forming a strict unitary matrix. Unlike LSTM and GRU, Axiom contains no forget gate — the unitary transition matrix guarantees lossless information preservation across arbitrary sequence lengths by mathematical construction. On standard long-range memory benchmarks, Axiom achieves 76.5–99.9% accuracy on the delayed copy task using 8,584 parameters, while LSTM (111,368 parameters, 13× more) scores 12.5–13.5% — random chance — across all delays on both GPU and TPU v6e. On the Adding Problem (Hochreiter & Schmidhuber, 1997), Axiom achieves MSE 0.00046 at length 200 versus LSTM's 0.00214. We further attach Axiom as a cross-chunk memory module to a frozen GPT-2 model and demonstrate that it retrieves facts from 7 chunks back (62.3% accuracy vs 8.3% baseline), with cluster analysis confirming genuine memory retrieval rather than pattern memorization. A controlled clean-vs-noisy experiment reveals the precise boundary of Axiom's advantage: on tasks requiring lossless recall of discrete signals, Axiom reaches 100% while LSTM reaches 42.8%; on tasks requiring noise filtering from continuous signals, LSTM's forget gate provides an advantage. We derive a closed form parallel forward pass reducing the unitary recurrence to a cumulative sum in a rotated eigenbasis, verified to within 2.81e 05 error. All results confirmed independently on NVIDIA GPU and Google TPU v6e. Technical Contributions: Householder Unitary Kernel: Implementation of a recurrence relation using a product of k Householder reflections to maintain strict orthogonality. XLA Optimization: A simplified single-sided rotation specifically designed to minimize the XLA graph size for TPU acceleration. Benchmark Dominance: Demonstration of solving the "Copy Task" at T=1000, a depth where traditional LSTMs and standard RNNs typically encounter total gradient collapse. Boundary characterization: controlled experiments identify exactly when unitary dynamics help (lossless discrete recall) vs when gating helps (noise filtering from continuous signals)

Read the paper · More papers on PaperTik