Investigating Effects of Variation in Evaluating Dialogue Quality Using Large Language Models
Amna Irum, Mirza Omer Beg · 2024
Evaluating dialogue systems is important to increase the capability and reliability of conversational models. Automatic metrics are needed because human evaluation tends to be expensive. Large language models (LLMs) have recently shown potential in evaluating various natural language generation (NLG) tasks. This work explores how well LLM evaluators assess open-domain dialogue tasks. We also investigated the sensitivity of LLMs in detecting and evaluating small changes in dialogue quality parameters. By introducing perturbations in dialogue responses, we aim to understand the sensitivity of LLM evaluators, which plays an important role in understanding the LLM-based evaluators’ systems.