A novel unsupervised fine-tuning method for text summarization, and highlighting the limitations of ROUGE score

Ala Alam Falaki, Robin Gras · Machine Learning with Applications · 2025

The limited availability of datasets for text summarization tasks and their similar characteristics (e.g. news articles) make it crucial to focus on unsupervised learning techniques to enable summarization across different domains. Moreover, since summarization produces text output, effective methods developed for news articles can be applied to other domains lacking sufficient labelled data. This study introduces a novel target selection process to be used as an unsupervised learning method for fine-tuning text summarization models with unlabeled data. The process involves two-steps: first, generating an extractive summary (Ext-Reference) from the article, and second, using an abstractive model to create a pool of candidate summaries. The most suitable summary (to be used as the target) is then selected by calculating the cosine similarity between the Ext-Reference’s embedding and each candidate’s embedding. Furthermore, this project underscores the limitations of the ROUGE score, which assigns a relatively low score to this method. However, extended analysis with various metrics, including using GPT-4 as a judge, demonstrates the effectiveness of this technique for fine-tuning models without a specific target reference. It highlights the importance of using a combination of metrics, like those included in the SumEvaluator package released alongside this paper. SumEvaluator package on Github: https://github.com/AlaFalaki/SumEvaluator .

Read the paper · More papers on PaperTik