Do We Really Need That Many Parameters In Transformer For Extractive Summarization? Discourse Can Help !

Wen Xiao, Patrick Huber, Giuseppe Carenini · 2020

The multi-head self-attention of popular transformer models is widely used within Natural Language Processing (NLP), including for the task of extractive summarization.With the goal of analyzing and pruning the parameterheavy self-attention mechanism, there are multiple approaches proposing more parameterlight self-attention alternatives.In this paper, we present a novel parameter-lean self-attention mechanism using discourse priors.Our new tree self-attention is based on document-level discourse information, extending the recently proposed "Synthesizer" framework with another lightweight alternative.We show empirical results that our tree self-attention approach achieves competitive ROUGE-scores on the task of extractive summarization.When compared to the original single-head transformer model, the tree attention approach reaches similar performance on both, EDU and sentence level, despite the significant reduction of parameters in the attention component.We further significantly outperform the 8-head transformer model on sentence level when applying a more balanced hyper-parameter setting, requiring an order of magnitude less parameters 1 .

Read the paper · More papers on PaperTik