A Context based Coverage Model for Abstractive Document Summarization

Heechan Kim, Soowon Lee · 2019

Automatic abstractive summarization is one of natural language processing fields that generates a sequence of words representing important information of the input document. The sequence-to-sequence model, which is widely used in abstractive summarization, has a repetition problem in which the same sub-pattern is repeatedly generated in decode phases. To solve this problem, various coverage models have been proposed in machine translation. In automatic summarization, unlike machine translation, the lengths of input and summary documents are very different. Because the summary document is a compressed form of important meaning of the input document. Due to the nature of automatic summarization, it is difficult to apply a word position-based coverage model in machine translation directly. For automatic summarization, we propose a context based coverage model to consider the coverage based on the compressed meaning of the input document. The context based coverage is defined as the accumulated weighted average of the encoded word meaning by the attention scores. This considers the meaning of words rather than the position of words in the input document. In the experiment with CNN/DailyMail dataset, the proposed model shows better performance than the previous researches.

Read the paper · More papers on PaperTik