RAGVAL: Automatic Dataset Creation and Evaluation for RAG Systems
Tristan Kenneweg, Philip Kenneweg, Barbara Hammer · 2024
Retrieval Augmented Generation (RAG) systems are widely used to augment Large Language Model (LLM) outputs with domain specific and time sensitive data. Each RAG system depends on a set of hyperparameters that strongly influence its performance. However, it is unclear how to evaluate different approaches on how to build and parameterize RAG setups for a given knowledge base. In this paper, we present a rigorous dataset creation and evaluation workflow that allows the quantitative comparison of different RAG strategies on a wide range of knowledge domains. Our method depends neither on human annotations nor handcrafted question sets and only requires the knowledge base the evaluated RAG system is given. We publish our code and exemplary dataset online.