Multi-Head Attention for End-to-End Neural Machine Translation

Ivan Fung, Brian Kan-Wing Mak · 2018

Inspired by the recent success of Google's Transformer model works have been done on borrowing the novel idea of multi-head attention to various applications under different architectures. Albeit latest works have adopted this idea using an end-to-end recurrent model on speech recognition and voice search, making use of a similar model on machine translation has not been attempted yet. In this work, we examine multi-head attention under the attention-based recurrent encoder-decoder frame-work, and conduct detailed analysis on the positional response of multiple heads. Through leveraging the essence of multi-head attention, we are capable of attaining a state-of-the-art result on IWSLT' 15 with 28.48 tokenized BLEU and 53.86% TER, which gives a 0.17 gain in BLEU and 0.37% reduction in TER. Similarly we achieve 25.58 tokenized BLEU and 55.03% TER on WMT'16, which provide a 0.40 gain in BLEU and 0.32% reduction in TER to the baseline model respectively. To the best of our knowledge, this is the first work1that evaluates the concept of multi-head attention in an end-to-end recurrent network on machine translation tasks.

Read the paper · More papers on PaperTik