$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan · arXiv (Cornell University) · 2024
A neutral, reproducible evaluation of the mcp-data-platform semantic knowledge layer. A single-shot four-arm ablation isolates cross-enrichment, search, and the memory/apply_knowledge lifecycle, and a cold-start study teaches facts one at a time and re-evaluates after each promotion. Every statistic is recomputed from raw run data committed under bench/results/ by the notebook bench/report/report.ipynb, with no network access and no API key. This archives the report together with the bench/results/ dataset it recomputes from.