Automatic Measurement of Dialogue Engagingness in Multilingual Settings
Amila Ferron · 2024
Expansive use of large language models (LLMs) as dialogue systems brings increased importance to the evaluation of the responses they generate. Although evaluation of qualities such as coherence and fluency are readily possible with well-established automatic metrics, engagingness is often measured with human evaluation -- a process that can be costly and slows the pace of development. Existing automatic metrics for engagingness have low to moderate correlation with human annotations, evaluate the response without the conversation history, are complicated to implement, or are designed for a specific dataset. Moreover, they have been tested exclusively on English conversations. Given that dialogue systems are increasingly available in languages beyond English, it is important to evaluate systems in more than one language. We propose that LLMs may be used for evaluation of engagingness in dialogue through prompting, and ask how prompt constructs compare in a multilingual setting. Our results give a prompt design taxonomy and indication of which strategies are the most effective. We find that using selected prompt constructs, including our comprehensive definition of engagingness, gives state-of-the-art performance on evaluation of engagingness in dialogue across multiple languages. We conclude that LLMs can be used for evaluation of engagingness in multiple languages through prompting alone.