Cheap Correlate or Discriminating Property? An Executable Validation Protocol for Consciousness- and Cognition-Theoretic Claims about AI
Patrick Butlin, Long, Robert, Eric Elmoznino, Yoshua Bengio, Jonathan Birch, Axel Constant, George E. Deane, Stephen M. Fleming, Chris Frith, Ji Xu, Ryota Kanai, Colin Klein, Grace W. Lindsay, Matthias Michel, Liad Mudrik, Megan A. K. Peters, Eric Schwitzgebel, Jonathan Steven Simon, Rufin VanRullen · arXiv (Cornell University) · 2023
Claims that an AI system "instantiates" a consciousness- or cognition-theoretic construct are often supported by a cheap correlate: aggregate dependence read as irreducible integration, ordinary shared-store efficiency read as global access, or a confidence-labelled signal read as metacognitive monitoring. Such inferences trade on a construct's appearance while omitting the property that carries the theory's content. The concern is not ours: the validation requirement (an indicator must be both sensitive and specific) is Seth and Bayne's (2022), the indicator-property programme is Butlin et al.'s (2023; 2026), and the proxy-can-be-gamed worry is Birch's (2022). We turn this desideratum into a prospective six-step protocol and audit how far one transparent chess-engine study actually satisfies it. Static feature dependence missed its specified support rule. A corrected 119-position temporal analysis missed its locally specified majority target, although a secondary post-hoc Bonferroni min-p test rejected the global no-departure null for both SI and Phi* (p = .0238). Shared-table configuration reduced wall time and total work, but this was an implementation effect and did not test a global-workspace construct. The internal signal provided no evidence of the specified directional ranking of later move persistence (AUROC .467), while its critic channel was nearly invariant, limiting the instrument. These are bounded case findings, not three independent datasets and not three definitive construct falsifications. The protocol therefore requires prospective theory-to-measure mapping, explicit nulls, sensitivity and specificity checks, power, and claim boundaries. Its generality remains a hypothesis until a second prospective substrate is studied. Status of this preprint. This is a working paper. The author's own gate for journal submission is HOLD/PIVOT: the protocol has not yet been applied to a second prospective substrate, and the theory-to-measure mapping has not yet been adjudicated by outcome-blind experts. It is deposited because the protocol (Section 4) and the reviewer checklist (Section 7.2) are usable independently of those two steps. Sections 1 to 4 and 7 are the proposed contribution; Section 5 is a bounded, single-substrate case audit.