Student Research Abstract: Evaluating Dialogue Summarization Using LLMs

Alison Wang · 2025

With the surge in audio data available today, there is a growing need for effective dialogue summarization. This study conducts two experiments using two LLMs, BART and Mistral, to assess dialogue summarization. The first experiment evaluates model performance, while the second examines the impact of upstream errors from Automatic Speech Recognition (ASR) and Machine Translation (MT) on summarization performance. Results indicate that SummaC, a commonly used evaluation metric, is unreliable for dialogue summarization. Additionally, Mistral's summarization performance is more sensitive to upstream errors than BART's.

Read the paper · More papers on PaperTik