ARLED: Leveraging LED-Based ARMAN Model for Abstractive Summarization of Persian Long Documents
Samira Zangooei, Amirhossein Darmani, Hossein Farahmand Nezhad, Laya Mahmoudi · 2025
With the increasing volume of textual data, reading and comprehending lengthy documents have become significant challenges, especially for researchers who need to extract useful information from academic articles. Automatic text summarization has emerged as a powerful tool for generating concise and informative summaries of long texts. Based on the applied approach, text summarization can be classified into extractive and abstractive methods. While extractive methods are more common due to their simplicity, they often fail to capture essential information. In contrast, abstractive summarization produces more coherent and informative summaries by understanding the underlying meaning of the text. In recent years, pre-trained models such as BERT, BART, and T5 have demonstrated remarkable advancements in abstractive summarization. However, the challenge of summarizing long documents remains, leading to the introduction of models like Longformer to overcome this limitation. This paper focuses on the abstractive summarization of Persian texts. A new dataset comprising 300,000 Persian academic papers from the Ensani website has been introduced. The ARMAN model, based on the Longformer architecture, is applied to generate summaries. Experimental results indicate the model's promising performance in Persian text summarization. This paper provides a comprehensive review of related works, presents the proposed methodology, analyzes the experimental results, and discusses potential directions for future research.