Research on Text Summary Generation Based on Bidirectional Encoder Representation from Transformers
Kai Wen, Zhou Lingyu · 2020
For Chinese automatic summarization, most of the generation methods are extractive, and the generative summary is not smooth, incoherent, and covers incomplete information. Compared with the traditional sequence-to-sequence model, Generative Adversarial Network (GAN) uses a reinforcement learning strategy The use of discriminator to guide generation has achieved good results in text generation. This paper proposes a pre-training method based on Bidirectional Encoder Representation from Transformers (BERT) and combined with LeakGAN model to generate abstracts. Firstly, using the bidirectional encoding characteristics of the BERT model, it can retain the original information well, and has a better effect when extracting features of words in the context to generate high-quality word vectors; secondly, for the current supervised generative model Both have the training problem of maximum likelihood estimation. This article uses the LeakGAN model that can decompose the task into different levels of sub-strategies, and uses hierarchical reinforcement learning to solve the characteristics of sparse rewards and generate a more accurate summary.