A Large-Scale Study of Relevance Assessments with Large Language Models Using UMBRELA
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Daniel Campos, Nick Craswell, Ian M. Soboroff, Jimmy Lin · 2025
There is substantial interest in applying large language models (LLMs) to provide relevance assessments in information retrieval (IR) applications from both industry and academia. To date, researchers and practitioners have presented several studies, but many questions remain. In this paper, we examine four different relevance assessment strategies: a fully manual process and three variants that rely on LLMs to different extents using our tool called UMBRELA. These were deployed in the TREC 2024 RAG Track on a diverse set of 77 runs from 19 teams in situ, which allowed us to correlate system rankings induced by the different approaches and to characterize tradeoffs between cost and quality. We find that system rankings produced by the three LLM-based strategies correlate well at the run level with those produced by fully manual assessments in terms of nDCG@20, nDCG@100, and Recall@100. On a topic-by-topic basis, the correlations are lower, and results using our setup indicate that increased human involvement does not improve correlations sufficiently to justify their costs. Our study suggests that LLMs can potentially replace fully manual judgments to measure run-level effectiveness in a coarse-grained manner.