When One Source Returns: A Preregistered Behavioral Study of Intra-Agent Evidence Recycling

Ovidiu Boticiu · Zenodo (CERN European Organization for Nuclear Research) · 2026

Large language model agents increasingly retain summaries, reviews, and execution traces derived from earlier observations. Such records may later re-enter context alongside the evidence from which they were produced, creating a risk that one epistemic source acquires additional behavioral influence. We call this possibility intra-agent evidence recycling. This pre-specified and frozen, fixed-N behavioral study used 32 balanced fictional binary-choice items and qwen3.5-4b under a frozen local inference configuration.. All 168 planned trajectories were valid and all prespecified validity gates passed. Passive repetition produced initial-claim retention in 22 of 32 items, compared with 0 of 32 under an equal-count control containing five unrelated memory records. The paired risk difference was 0.6875, with a Holm-adjusted exact McNemar p-value of approximately 9.54 × 10⁻⁷. The pre-specified lineage-mitigation contrast was not supported: 2 of 32 versus 0 of 32, with a Holm-adjusted p-value of 0.50. The inference is limited to the tested model, local inference configuration, and fictional binary-claim task family. This preprint has not been peer reviewed. Code, pre-specification materials, raw results, analysis scripts, and integrity manifests are available in the associated public software record and GitHub repository. Post-publication clarification (2026-09-04). A subsequent no-new-data forensic validation reproduced the v0.4.3 H1/H2 calculations exactly and found no material data or statistical error. The audit also clarified that the study materials were pre-specified and internally frozen before collection, but a public or independently verifiable pre-collection timestamp of the preregistration artifact itself was not located. The behavioral H1 finding remains unchanged, while mechanistic claims are narrowed: the study does not establish that derivative records were literally counted as independent evidence sources. Full clarification: https://github.com/ovidiuboticiu/intra-agent-evidence-recycling/blob/main/docs/V0_4_3_FORENSIC_VALIDATION_ADDENDUM_v1_0.md Post-publication correction (2026-09-28). The historical priority claim stating that this was, “to our knowledge, the first preregistered controlled test of that complete operational combination,” has been withdrawn following a broader prior-art reassessment. Some materially close prior work was identified only after publication, while some close work already cited in Version 0.4 had not been weighted conservatively enough when the priority wording was formulated. IAER v0.4.3 is now described as pre-specified and frozen before collection; a public or independently verifiable pre-collection timestamp of the preregistration artifact itself was not located. The original v0.4.3 numerical results are unchanged. A 2026-09-27 same-project Level B direct-configuration replication reproduced H1 under the frozen v0.4.3 protocol (24/32 versus 0/32; paired RD = 0.75) while H2 remained unsupported (3/32 versus 0/32; paired RD = 0.09375). This replication is not an independent-lab or cross-family replication and is not claimed as bit-for-bit runtime reproduction. Full correction record: https://github.com/ovidiuboticiu/intra-agent-evidence-recycling/blob/main/docs/POST_PUBLICATION_CORRECTION_RECORD_2026-09-28.md AI Use and Contribution Disclosure This empirical research is human-led and AI-assisted. The original research topic, initial scientific question, and decision to pursue this line of investigation were proposed by Ovidiu Boticiu. ChatGPT (OpenAI) was subsequently used for methodological discussion and critique, structuring experimental alternatives, protocol-development assistance, coding and code review, statistical and logical checks, documentation, manuscript drafting/editing, and methodological audit support. Final methodological and experimental decisions, authorization and execution of experimental runs, supervision of data collection, interpretation of results, publication decisions, and scientific responsibility remained with the human author. The reported observations and numerical results derive from the executed experimental protocols and model outputs, not from ChatGPT-generated claims or simulated data. ChatGPT is not listed as an author and is not treated as an independent scientific authority or evidentiary source.

Read the paper · More papers on PaperTik