BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models

Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, Iryna Gurevych · arXiv (Cornell University) · 2021

A neutral, reproducible evaluation of what hyperlinks between the pages of an organizational wiki contribute to a tool-using LLM agent, measured on mcp-data-platform, an open-source MCP (Model Context Protocol) server that exposes stored knowledge pages to agents through search and fetch tools. A pre-registered two-arm contrast plants one deterministic generated corpus twice at 50, 500, and 5000 pages — once with its cross-references as real, followable links, once with each link rewritten as an ordinary sentence carrying the same meaning — and measures grounded coverage, search cost, fetch provenance, and elicited completeness claims across 99 episodes of an Anthropic Claude agent writing operational documents whose required facts are spread across the corpus. The headline is the study's pre-registered instrument kill, presented as a finding: facts certified unreachable by two independent instruments for every task-derived phrasing were recovered by the link-stripped arm through read-derived queries (the agent reads a prose mention and re-searches in the named institution's own vocabulary), so "unreachable by search" must hold against read-derived queries, which meaning-preserving prose cannot provide. What the archives affirmatively show: grounded coverage at ceiling in every cell (discovery is not enumeration; roughly eleven page reads ground every fact against a 5000-page store), a cost separation that grows with scale (searches per grounded fact flat for the linked arm across two orders of magnitude, roughly doubling without links), and full-depth link traversal when search is unavailable, where the stripped arm has no route at all. Every statistic is recomputed from raw run data committed under bench/results/ by bench/reports/graph-completion/graph_tables.py, offline, with no network access and no API key. This archives the report (docs/reference/benchmark-report-graph-completion.md) together with the run-data snapshot it recomputes from.

Read the paper · More papers on PaperTik