AraBART: a Pretrained Arabic Sequence-to-Sequence Model for Abstractive Summarization
Moussa Kamal Eddine, Nadi Tomeh, Nizar Y. Habash, Joseph Le Roux, Michalis Vazirgiannis · 2022
Like most natural language understanding and generation tasks, state-of-the-art models for summarization are transformer-based sequenceto-sequence architectures that are pretrained on large corpora.While most existing models focus on English, Arabic remains understudied.In this paper we propose AraBART, the first Arabic model in which the encoder and the decoder are pretrained end-to-end, based on BART (Lewis et al., 2020).We show that AraBART achieves the best performance on multiple abstractive summarization datasets, outperforming strong baselines including a pretrained Arabic BERT-based model, multilingual BART, Arabic T5, and a multilingual T5 model.AraBART is publicly available on github 1 and the Hugging Face model hub 2 .