Data-to-text Generation with Pointer-Generator Networks
Mengzhu Liu, Zhaonan Mu, Jieping Sun, Cheng Wang · 2020
According to attention mechanism-based encoder-decoder model for data-to-text generation does not fully consider the structure information of the slot-value pair data and the integrity of the expression. In this paper, firstly, the pointer-generator network is used to copy the word from the input to reference the slot-value pairs data and handle the problem of the out of vocabulary words and rare words. Then propose the slot-attention mechanism to calculate the attention score using the attribute sequence and value sequence simultaneously, construct the context vector separately for the attribute sequence and the value sequence, and input them into the decoder to alleviate the problem of assigning values to the wrong attributes. Secondly, this paper introduces the coverage mechanism to use the historical attention information to calculate the attention score so that the model considers more about the unexpressed attributes to alleviate the situation where some attributes appear repeatedly while the others do not appear in the generated text. Finally, the slot-value pair data are content words, which are not aligned one-to-one with the generated text. This paper uses the attention distribution gate to dynamically control the softness of the attention distribution. When generating a content word, the model focus on the most relevant one word to capture the semantic information, attention distribution becomes more harder; When generating a function word, the model focus on the most relevant several words to capture the syntactic information, and the attention distribution becomes smoother. Experimental results on the E2E dataset demonstrate that the proposed model can further improve the quality of the generated text.