Improving Fully Non-Autoregressive Translation with Pre-Trained Language Models
Zijian Fu, Shuheng Wang · 2024
In recent years, Non-Autoregressive Translation (NAT) has received lots of attention because of its outstanding decoding speed. However, there is a certain gap between the NAT model and its autoregressive comparator. It, how to effectively narrow this gap, remains to be further investigated. In this work, we use BERT to improve the performance of NAT model. Specifically, we add a CTC model on the top of BERT for extracting semantic and learning alignment. In order to test the effect of the model, we conducted the experiments on three public translation datasets. The experimental results show that our model can significantly outperforms existing non-autoregressive baselines.