SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization
Mathieu Ravaut, Shafiq Joty, Nancy F. Chen · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) · 2022
Sequence-to-sequence neural networks have recently achieved great success in abstractive summarization, especially through fine-tuning large pre-trained language models on the downstream dataset.These models are typically decoded with beam search to generate a unique summary.However, the search space is very large, and with the exposure bias, such decoding is not optimal.In this paper, we show that it is possible to directly train a secondstage model performing re-ranking on a set of summary candidates.Our mixture-of-experts SummaReranker learns to select a better candidate and consistently improves the performance of the base model.With a base PEGASUS, we push ROUGE scores by 5.44% on CNN-DailyMail (47.16 ROUGE-1), 1.31% on XSum (48.12 ROUGE-1) and 9.34% on Reddit TIFU (29.83 ROUGE-1), reaching a new state-of-theart.Our code and checkpoints will be available at https://github.com/ntunlp/ SummaReranker.