Complementary Representation of ALBERT for Text Summarization
Wenying Guo · Proceedings/Proceedings of the ... International Conference on Software Engineering and Knowledge Engineering · 2021
Pretraining has proved to be an effective strategy to learn the parameters of the deep neural network.It captures the world knowledge that can be adapted to downstream tasks.Text summarization based on ALBERT [1] outperformed previous work by a large margin.However, they only use the final layer as a contextualized representation of the input text.Multiple studies have proven that intermediate layers also encode the rich hierarchy of linguistic information.In this paper, we propose a Fast Complementary Representation Network (FCRN), which dynamically incorporates linguistic knowledge spread across the entire ALBERT for extractive selection.Different from previous work, we measure the importance of hidden layers by all sentence representations rather than all token embeddings, which can filter nonsignificant words and takes six times less time during training.FCRN first obtains the importance of each layer by sentence embeddings and then automatically absorbs the supplementary information to ALBERT's output.We conduct experiments on CNN/DailyMail and XSum datasets.The results show that our model obtains higher ROUGE scores.