Domain-Specificity of Refusal Representations in Large Language Models
Charlie Tucker, Maikel León · Proceedings of the ... International Florida Artificial Intelligence Research Society Conference · 2026
Modern Large Language Models are trained using a variety of techniques to reject prompts that lead to harmful output. Recent work has shown that a model's likelihood of refusing a prompt is mediated by a single direction in its activation subspace. We investigate the domain specificity of refusal representations by extracting and comparing refusal directions across distinct knowledge domains, and analyzing how activation magnitude along these directions varies with prompt category. These experiments give insight into how models learn to refuse during post-training. By demonstrating the universality of this refusal direction, we highlight a systemic vulnerability: removing a single geometric feature compromises safety guardrails globally, across all distinct knowledge domains.