Unleashing the Potential of Attention Model for News Headline Generation
Yong Liao, Kui Meng, Jianshen Zhang, Gongshen Liu · 2020
Headline generation is a special summarization generation task and the difficulty lies in requiring the generated headline to be concise, fluent and informative. Limited by the ability of commonly used encoder and decoder modules to capture long-term dependencies in seq2seq tasks, previous work rarely researched headline generation by end-to-end methods. However, the recent success of Transformer model and its subsequent improved versions demonstrate their remarkable performance on seq2seq tasks, which provide us with a feasible solution. In this paper, we propose a novel model Transformer(XL)-CC to generate headline from the perspective of understanding the whole text, the segment-level recurrence mechanism and relative positional encoding make our model learn ultra-long-term dependencies. In addition, we combine the copy and coverage mechanisms to generate more readable titles. Experimental results on the NYT and Chinese LSCC news datasets also confirm that our method significantly achieves better performance on the headline generation task.