A Pre-Trained Language Model Based on LED for Tibetan Long Text Summarization
Xinpeng Ouyang, Xiaodong Yan, Minghui Hao · 2024
As one of the important tasks in the field of natural language processing, text summarization aims to obtain important information from a large number of text data and extract the main content of the text. Text summarization can be categorized into long and short text summarization based on the length of the text. According to the literature search, there is no research on Tibetan long text summarization and no publicly available dataset of Tibetan long text summarization. To promote the development of Tibetan long text summarization research, this paper first constructed a Tibetan long text summarization dataset containing nearly 12000 samples. To process Tibetan long text, we designed a model called Ti-LED for Tibetan long text summarization task based on the LED structure and pre-trained it using the "two-phase multi-task" strategy. The maximum coding sequence length of the Ti-LED model reaches 4096, which can effectively solve the problem that the coding sequence length of other models is too short. The experimental results in multiple datasets and comparison of multiple models show that the TiLED model built in this paper has a strong summary generation capability, which is superior to other existing models.