Greek Wikipedia: A Study on Abstractive Summarization

Nikolaos Giarelis, Charalampos Mastrokostas, Nikos I. Karacapilidis · 2024

In the research area of Natural Language Processing (NLP), text summarization (TS) has been defined as the automatic composition of a cohesive and articulate summary, which encapsulates the main ideas and themes from a single or multiple documents. Recent developments from the areas of Deep Learning (DL) and Natural Language Understanding (NLU) have facilitated the development of abstractive TS transformer-based models, which demonstrate better performance than classical extractive ones. In any case, the vast majority of current NLP research concentrates on high-resource languages (e.g., English). Dealing with modern Greek, this paper introduces and elaborates a novel abstractive TS dataset comprising 93,433 Greek Wikipedia articles and their summaries assigned by human editors. The paper also proposes a series of DL abstractive TS models that were fine-tuned on this dataset for the task of Greek article summarization. A thorough experimentation was conducted for the comparative assessment of the proposed models against well-known extractive summarization ones, using the test subset of the dataset. The results reveal that the proposed abstractive models outperform the extractive ones across various TS evaluation metrics. To enhance the reproducibility of our work, we make publicly available the corresponding dataset, our best performing model and the experimentation code.

Read the paper · More papers on PaperTik