An Extraction-Abstraction Hybrid Approach for Long Document Summarization

Huang Si, Rui Wang, Qing Jie Xie, Lin Li, Yongjian Liu · 2019

In this paper, we propose a hybrid model of extractive and abstractive methods to tackle the long document automatic summarization task. The model first trains an extractor to extract salient sentences from the original text. Next, these salient sentences are put together to get a condensed version of the original text. Then we use the abstractive model to rewrite the extracted sentences to get the final summary. In order to avoid the exposure bias, reinforcement training is used to optimize the proposed model. Experiments in NLPCC2017 Shared Task 3 show that our models achieve competitive performance. Additionally, the ROUGE score of our model exceeds the score of the state-of-the-art model in the original NLPCC2017 Shared Task 3, where a sentence summary is generated from each Chinese news article.

Read the paper · More papers on PaperTik