Concept-injection introspection under controls: a replication on Qwen2.5

Theia Pearson-Vogel, Martin Vanek, Raymond Douglas, Jan Kulveit · arXiv (Cornell University) · 2026

We replicate a published concept-injection "introspection" result on Qwen2.5-Coder-32B-Instruct. The replication succeeds on its own terms: control-trial P(yes) 0.16% vs ~0.2% reported; injection effect 0.74 nats [95% CI 0.28, 1.14] vs ~0.4 reported; a late-layer logit-lens peak that attenuates at the output. Adding two controls absent from the original protocol changes the reading: a same-norm random vector raises "yes" at least as much as the concept vector, the mean of all concept vectors raises it more, and the injected concept is named in 0 of 250 trials. On Qwen2.5-7B-Instruct, concept vectors shift detection answers beyond random vectors (S = +4.90 nats [3.01, 6.70]), but a single mean vector does the same: the model responds to "conceptness", not concept identity. Steering an "exhausted / full of energy" direction does not yield a coherent numeric self-report; a reverse-keyed item and an unrelated direction show why. All confirmatory analyses were preregistered and run once on held-out data; deviations are listed.

Read the paper · More papers on PaperTik