A Neural Machine Translation Model Based on Sequence to Dependency

Ting Yang, Shinan Zhao, He Chen, Ting Liu · Journal of Physics Conference Series · 2021

A supervised self-attention network is proposed to introduce the dependency structure into the Transformer model. This scheme is mainly proposed based on the characteristics of the self-attention mechanism in Transformer, which converts the dependency syntax tree into two equivalent adjacency matrices, and then uses the adjacency matrix to supervise the self-attention network in the Transformer's encoder. So that Transformer learns how to model the dependency structure of the source language. Although this scheme is simple, the effect is remarkable. In particular, there is no need to modify the network structure, as long as two additional loss functions are added in the process of training Transformer. In the decoding process, there is no need to use an external syntactic analyzer to perform syntactic analysis on the source language, but it can automatically construct the dependent syntactic structure of the source language. Experiments demonstrate that the quality of the constructed dependency syntactic structure is also good.

Read the paper · More papers on PaperTik