Multi-document summarization algorithm based on significance sentences

Liu Na, Ying Lü, Xiaojun Tang, Wang Hai-wen, Xiao Yan Peng, Ming-Xia Li · 2016

Latent Dirichlet Allocation (LDA) has been used to generate text corpora topics recently. The basic idea of most LDA is that documents are represented as random mixtures over latent topics, each topic is characterized by a distribution over words. However, the main task of multi-document summarization is sentences selection. For generic multi-document summarization, we propose a multi-document summarization algorithm based on significance sentences. Firstly, our method proposes a sentence_LDA model which represents topics as a mixture of sentences not words. Secondly, our proposed method introduces three different criteria to determine the significance of sentences. Multi-document summarization is consists of significance sentences. The experiments showed that the proposed algorithm achieved better performance compared the other state-of-the-art algorithms on DUC2002corpus.

Read the paper · More papers on PaperTik