TGSum: Build Tweet Guided Multi-Document Summarization Dataset
Ziqiang Cao, Chengyao Chen, Wenjie Li, Sujian Li, Furu Wei, Ming Zhou · WORLD SCIENTIFIC eBooks · 2017
The development of summarization research has been significantly hampered by the costly acquisition of ref-erence summaries. This paper proposes an effective way to automatically collect large scales of news-related multi-document summaries with reference to social me-dia’s reactions. We utilize two types of social labels in tweets, i.e., hashtags and hyper-links. Hashtags are used to cluster documents into different topic sets. Also, a tweet with a hyper-link often highlights certain key points of the corresponding document. We synthesize a linked document cluster to form a reference sum-mary which can cover most key points. To this aim, we adopt the ROUGE metrics to measure the cover-age ratio, and develop an Integer Linear Programming solution to discover the sentence set reaching the up-per bound of ROUGE. Since we allow summary sen-tences to be selected from both documents and high-quality tweets, the generated reference summaries could be abstractive. Both informativeness and readability of the collected summaries are verified by manual judg-ment. In addition, we train a Support Vector Regression summarizer on DUC generic multi-document summa-rization benchmarks. With the collected data as extra training resource, the performance of the summarizer improves a lot on all the test sets. We release this dataset for further research1.