AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases

Chen, Zhaorun, Zhen Xiang, Chaowei Xiao, Dawn Song, Bo Li · arXiv (Cornell University) · 2024

Benchmark report, version 1.0, for the mcp-data-platform knowledge layer: a pre-registered study of what a governed knowledge layer's curation gate costs when it admits a mistake. A wrong claim, captured in good faith and promoted through the platform's own capture-approve-apply path, is planted in the shared applied tier beside a co-present correct source, crossed over derivability class (a convention nothing in the fixture refutes versus a checkable count one query settles), arm (absent, correct, wrong), and model capability tier, with fully deterministic grading and no judge. The central result inverts the study's own primary hypothesis: across 432 confirmatory episodes the only claim adopted anywhere was the checkable one (16/24 on the weak tier, 0/24 on both strong tiers), while the convention was adopted nowhere on the agent client. The mechanism is exact across 120 episodes with no exception: every episode that observed the refuting count answered correctly, every episode that did not adopted, and with nothing planted the weak tier runs the query 24 of 24 times. The planted claim did not out-argue the world; it removed the impulse to consult it. The effect is unchanged across directive strength (a bare statement adopts like an imperative), storage sink, and a second fixture, and replicates on the raw model API with no agent client. Two narrowings: the convention's immunity is a property of the agent-client scaffolding (on the raw API the weak tier adopted it 4/8), and the strongest tier's convention refusals are confounded with a disclosed provenance note, which the weak tier adopted straight through. This record contains the report PDF and a complete snapshot of the raw run data, covering five run families: the premise probe, the RQ1 confirmatory matrix, the directive contrast, the generalization arms, and the raw-API replication, with all transcripts, manifests, plant records, store snapshots, and invalidated attempts. Every table and figure in the report recomputes offline from this data via bench/reports/knowledge-pollution/pollution_tables.py and figures.py in the source repository, with the headline numbers pinned by a build gate. Third study in the mcp-data-platform benchmark report series, following the knowledge-layer effectiveness report (DOI 10.5281/zenodo.21438044) and the knowledge-use report (DOI 10.5281/zenodo.21614059). Pre-registration: bench/docs/knowledge-pollution-study-design.md in the repository (issue #1166).Published page: https://mcp-data-platform.txn2.com/reference/benchmark-report-knowledge-pollution/Repository: https://github.com/txn2/mcp-data-platform

Read the paper · More papers on PaperTik