When Language Becomes Workflow: Provenance Failures in Agentic LLM Systems
Anuar Kiryataim Contreras Malagón · Zenodo (CERN European Organization for Nuclear Research) · 2026
Tool-bearing LLM agents are often evaluated as if the final assistant response were the primary safety artifact. This paper argues that such evaluation is insufficient for multi-agent and tool-mediated systems, where user language can be routed, summarized, delegated, transformed into tool arguments, written into tickets, used as policy-like evidence, or passed into human handoff. The paper develops a taxonomy of provenance and workflow failures in agentic LLM systems, identifying five primary families: action-layer inconsistency, handoff laundering, identity and authorization drift, policy-provenance failure, and harmful-purpose reclassification.