Decoupling Privacy Noise from Optimization in Transformer Forecasting
Bhagiradh Kantheti, Carlos A. Paz de Araújo · Machine Learning and Knowledge Extraction · 2026
Strong differential privacy often collapses utility in transformer-based time-series forecasting because noise is injected directly into high-dimensional gradients (e.g., DP-SGD), severely corrupting the optimization process. We introduce Low-Dimensional Feature-Path Privacy for Transformers (LDPT), which enforces privacy by routing calibrated perturbations through a low-dimensional feature bottleneck (D=16) that is independent of the model parameter count. LDPT implements noise via classically simulated quantum channels (Lindblad/depolarizing dynamics) and finite-shot POVM measurements, providing an auditable mapping from privacy budget ε to perturbation magnitude while keeping the transformer gradients clean. Across the ETT datasets and multiple prediction horizons, LDPT substantially preserves forecasting utility under its native local ε-QDP guarantee. At a nominal per-pass ε=0.1, LDPT limits MSE degradation to under 6%. In contrast, DP-SGD with global (ε,δ)-DP applied to the identical transformer architecture suffers over 100% MSE degradation. Because these methods operate under different privacy definitions (local ε-QDP vs. global (ε,δ)-DP), this comparison illustrates the impact of noise placement rather than equivalent privacy protection. To isolate the effect of the calibration mechanism, we further evaluate a classical Gaussian mechanism on the same feature-path bottleneck, which requires orders-of-magnitude larger noise and severely degrades utility. Membership inference attacks confirm that LDPT does not amplify membership leakage beyond the non-private baseline. These results demonstrate that decoupling privacy noise from optimization through low-dimensional feature-path placement and tight channel-based calibration is critical for practical privacy-preserving transformer forecasting.