A Simple but Effective Way to Improve the Performance of RNN-Based Encoder in Neural Machine Translation Task
Yan Liu, Dongwen Zhang, Lei Du, Zhaoquan Gu, Jing Qiu, Qingji Tan · 2019
Armed with attention mechanism, the recurrent neural network-based encoder-decoder model (or sequence to sequence model) has become the standard architecture to tackle many sequence Nature Language Processing (NLP) tasks. With regard to the neural machine translation (NMT) tasks, our paper proposed a new architecture to proficiently mines the ability of attention mechanism and stacked recurrent neural networks. As a lot of work has given proved that each layer of the stacked recurrent neural networks learns different aspects of a sequence. That means the information represented by each layer is important in terms of the translation task. However, usually, most work just simply adopt stacked recurrent neural networks as a whole part as the encoder or decoder and then combined them with the attention mechanism. While our work creatively uses the attention mechanism to explore each layer of the encoder. In this way, many aspects of the sequence could be learned. For example, linguistic features and semantics information of a word in a sentence could be clearly captured and finally influences the generation of the current translation word. Experiments have shown the effectiveness of our model and an average of 5.67 points BLEU scores were promoted.