RESA: Relation Enhanced Self-Attention for Low-Resource Neural Machine Translation
Xing Hui Wu, Shumin Shi, Heyan Huang · 2021
Transformer-based Neural Machine Translation models have achieved impressive results on many translation tasks. In the meanwhile, some studies prove that extending syntax information can be explicitly incorporated to provide further improvements especially for some low-resource languages. In this paper, we propose RESA: the relation enhanced self-attention for Transformer which can integrate source side dependency syntax. More specifically, dependency parsing produces two kinds of information: dependency heads and relation labels, compared to the previous works only pay attention to dependency heads information, RESA use two methods to integrate relation labels as well: 1) Hard-way that uses a hyper parameter to control the information percentage after mapping relation labels sequence to continuous representations; 2) Gate-way that employs a gate mechanism to mix word information and relation labels information. We evaluate our methods on low-resource Chinese-Tibetan and Chinese-Mongol translation tasks, and the preliminary experimental results show that the proposed model achieves 0.93 and 0.68 BLEU scores gain compared to the baseline model.