A Lightweight Transformer Model with Dynamic Sparse Mask for Neural Machine Translation

Nastaran Asadi, Babak Golbabaei, Yirong Kan, Renyuan Zhang, Yasuhiko Nakashima · 2025

Neural machine translation (NMT) models, especially Transformers, have high translation quality but suffer from scalability issues due to the quadratic complexity of the attention mechanism. This drawback leads to high memory usage and computational overhead, making large-scale translation inefficient. In addition, maintaining translation quality while reducing computational costs remains a key challenge. To alleviate these issues, this work presents dynamic sparse masks for Transformers, which fine-grained control sparsity based on percentile thresholds. Our method greatly reduces computation amount in Transformer inference process while preserving translation accuracy by selectively retaining the most significant attention scores. Furthermore, learning the curriculum is incorporated to improve model accuracy by gradually structuring the training process. Experiments are conducted on the WMT2014 EnglishFrench dataset to verify the effectiveness of our model on NMT tasks. Compared to the standard Transformer model, our model achieves an improvement of 19.38 in the Bilingual Evaluation Understudy (BLEU) score and reduces 31.4% multiply-accumulate (MAC) operations in Transformer inference process, which is crucial to reducing power consumption in hardware.

Read the paper · More papers on PaperTik