Team NLLG submission for Eval4NLP 2023 Shared Task: Retrieval-Augmented In-Context Learning for NLG Evaluation
Daniil Larionov, Vasiliy Viskov, George Kokush, Alexander Panchenko, Steffen Eger · 2023
In this paper, we introduce a novel approach for evaluating natural language generation (NLG) using retrieval-augmented in-context learning.Our method empowers practitioners to leverage large language models (LLMs) for diverse NLG evaluation tasks without the need for finetuning.We put our approach to the test in the context of the Eval4NLP 2023 Shared Task, specifically in translation evaluation and summarization evaluation subtasks.The results indicate that retrieval-augmented in-context learning holds great promise for the development of LLM-based NLG evaluation metrics.Future research directions involve investigating the performance of various publicly available LLM models and identifying the specific LLM attributes that contribute to enhancing metric quality.