A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment

Negar Arabzadeh, Charles L. A. Clarke · 2025

Large Language Models (LLMs) are increasingly used to automate relevance judgments for information retrieval (IR) tasks, often demonstrating agreement with human labels that approaches inter-human agreement. To assess the robustness and reliability of LLM-based relevance judgments, we systematically investigate impact of prompt sensitivity on the task. We collected prompts for relevance assessment from 15 human experts and 15 LLMs across three tasks-binary, graded, and pairwise-yielding 90 prompts in total. We compare LLM-generated labels with TREC official human labels using Cohen's κ and pairwise agreement measures. In addition, we compare human- and LLM-generated prompts and analyze differences among different LLMs as judges. We release all data and prompts at https://github.com/Narabzad/prompt-sensitivity-relevance-judgements/.

Read the paper · More papers on PaperTik