Dynamic Skill Safety Filters via Differential Failure Attribution for Autonomous Robot Agents

Shivam Johri · Zenodo (CERN European Organization for Nuclear Research) · 2026

Agentic robots increasingly act by composing a library of skills— parameterized subroutines such as grasp, pour, heat, or cut—selected on the fly by a large-language-model (LLM) policy. Recent evidence shows that individual skills can be harmful: in the wrong context a nominally useful skill induces a safety violation or silently degrades the task. Prevailing mitigations are either a static blocklist that bans ``dangerous'' skills outright—destroying task success because those skills are genuinely needed—or trusting the LLM to police itself, which is unreliable under distribution shift and adversarial pressure. We propose a Dynamic Skill Safety Filter (DSSF): a lightweight, model-agnostic runtime layer that maintains an online, context-conditioned estimate of each skill's failure propensity via differential failure attribution, and disables a skill only in contexts where it is currently estimated to be risky—prompting the agent to first mitigate the hazard and then act. On a symbolic long-horizon manipulation benchmark with a hidden, context-dependent, stochastic harm model, and with skill selection driven by a free open-weight LLM, DSSF cuts safety violations from 0.70 to 0.00 per episode (-100%) while preserving full task success (1.00). A static blocklist eliminates violations only by collapsing success to 0.00 and wrongly disabling safe skills 85% of the time. DSSF matches an oracle that knows the true harm model on success, violations, harmful-skill activations, and regret, while cutting the false-disable rate to 0.24. The method is pure software, needs no GPU, robot, or additional data, and adds negligible latency.

Read the paper · More papers on PaperTik