FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation
Farima Fatahi Bayat, Lechen Zhang, Sheza Munir, Lu Wang · 2025
The rapid adoption of language models (LMs) across diverse applications has raised concerns about their factuality, i.e., their consistency with real-world facts.We introduce VERIFY, an evidence-based evaluation pipeline that measures LMs' factuality in real-world user interactions.VERIFY considers the verifiability of LM-generated content and categorizes content units as Supported, Unsupported, or Undecidable based on Web-retrieved evidence.Importantly, factuality judgment by VERIFY more strongly correlates with human evaluations than existing methods.Using VER-IFY, we identify "hallucination prompts," i.e., those that frequently elicit factual errors in LM responses.These prompts form FACTBENCH, a dataset of 1K prompts spanning 150 topics and tiered into Easy, Moderate, and Hard prompts.We benchmark widely-used openweight and proprietary LMs from six families, yielding three key findings: (i) LMs' factual precision declines from Easy to Hard prompts, (ii) factuality does not necessarily improve with scale; Llama3.1-405B-Instructperforms comparably to or worse than its 70B variant, and (iii) Gemini1.5-Proshows a notably higher refusal rate, with over-refusal in 25% of cases.