Error Analysis of using BART for Multi-Document Summarization: A Study for English and German Language.
Timo Johner, Abhik Jana, Chris Biemann · DSpace repository (University of Tartu) · 2021
Recent research using pre-trained language models for multi-document summarization tasks lacks a deep investigation of potential erroneous cases and their possible application in other languages.In this work, we apply a pre-trained language model (BART) for multi-document summarization (MDS) task, both with fine-tuning and without finetuning.We use two English datasets and one German dataset for this study.First, we reproduce the multi-document summaries for the English language by following one of the recent studies.Next, we show the applicability of the model to the German language by achieving state-of-the-art performance on German MDS.We perform an in-depth error analysis of the followed approach for both languages, which leads us to identify the most notable errors, from made-up facts to topic delimitation.Lastly, we quantify the amount of extractiveness.