Improving Machine Translation With PPO and Token-Adaptive Early Exit in Transformer-Based Models
Ronit Gangwani, Diksha Kumaraguru, Ananya C. Sawant, A. Devipriya · IEEE Access · 2026
Transformer-based machine translation (MT) systems such as MarianMT achieve high translation quality but suffer from high inference latency due to autoregressive decoding, where each token is processed through all decoder layers. This limits their use in real-time and resource-constrained settings. A token-adaptive early-exit framework is introduced to reduce inference cost by allowing tokens to terminate decoding at intermediate layers when confidence estimates exceed a learned threshold. Lightweight exit heads are attached to each decoder layer to produce token logits and confidence scores, while Proximal Policy Optimization (PPO) is employed during training to regularize confidence estimation and guide exit decisions. The full latency–quality Pareto frontier is characterized across three translation benchmarks—WMT16 English–Romanian (En–Ro), WMT16 English–German (En–De), and WMT14 English–French (En–Fr)—spanning four operating points (λ2∈ {0.05, 0.10, 0.20, 0.40}). At the optimal operating point (2=0.05), the proposed model achieves 96% baseline translation quality (BLEU scores for En-Ro, 61.4 and 63.9) at a 1.21 × increase in processing speed and a 95 th -percentile latency decrease from 51.4 ms / token to 42.6 ms / token . The full range of latency-quality trade-offs span speedup values of 1.21-1.72 × , for practitioners with tighter latency constraints. Additionally, consistent depth clustering for all examined language-pairs (exit-depth = 3.3 - 3.5 layers) suggest that the learned exit-policy is capturing generalizable syntactic difficulty characteristics and not solely based on corpus behavior. These results provide insight into the latency-quality trade-off and show adaptive, confidence-guided depth-allocation as an efficient layer to improve performance for latency-sensitive MT applications.