Pencils Down! Automatic Rubric-based Evaluation of Retrieve/Generate Systems
Naghmeh Farzi, Laura Dietz · 2024
Current IR evaluation paradigms are challenged by large language models (LLMs) and retrieval-augmented generation (RAG) methods. Furthermore, evaluation either resorts to expensive human judgments or lead to an over-reliance on LLMs.