Improving Word Alignment by Adding Gromov-Wasserstein into Attention Neural Network

Yan Huang, Tianyuan Zhang, Huidong Zhu · Journal of Physics Conference Series · 2022

Abstract Statistical machine translation systems usually break the translation task into two or more subtasks and an important one is finding word alignments over a parallel sentence bilingual corpus. We address the problem of introducing word alignment for language pairs by developing a novel neural network model that can applied to other generative alignment models. We use Multi-layer attention model and multi-layer model with multi-head-attention mechanism on each layer provides superior translation quality. It can be trained on bilingual data without relying on word alignment. In this paper, we cast the correspondence problem directly as an optimal distance problem. We use the Gromov-Wasserstein distance to calculated how similarities between word pairs are related across languages. The resulting alignments dramatically outperform the GIZA++ and FastAlign approach, these alignments are comparable on public data sets.

Read the paper · More papers on PaperTik