Reconstructing Review-Mediated Change from Software Archives: An Instrumentation Study Using CROP
James Michael Barlow · Zenodo (CERN European Organization for Nuclear Research) · 2026
A technical note and reproducible research object continuing Fregnan, Petrulio and Bacchelli's work on changes made during code review. It asks whether language-model-assisted reconstruction can establish a stricter experimental unit: an action absent before review, requested by a reviewer, and implemented in the next revision. Five apparatus versions applied to the Code Review Open Platform (CROP) exposed persistence, moving-parent, semantic-exactness, multi-action and machine-interface errors. The final source-qualification fixtures failed, so the sealed fresh sample was not opened and no comparative effect of an explicit representation of organisational judgement was estimated. A subsequent calibration deterministically joined 164 cases from Fregnan et al.'s released inter-rater material to their earlier CROP discussions. On 144cases of initial human agreement, Qwen3-32B reproduced the retained binary distinction under true review history while change-only andmatched-other-history controls remained at chance balanced accuracy. A revised frozen run using immutable comment identifiers achieved 92.1% true-historybalanced accuracy and passed all six advancement checks. This is a bounded computational reproduction of the released coding distinction, not anindependent replication of its population prevalence and not validation of exact review-mediated action reconstruction. The note concludes that language models can assist archival coding but do not by themselves resolve its metrology, and proposes a staged successor designusing deterministic Git normalisation, auditable coding, independent admission and a measured source-yield pilot. Retained successful hosted runs consumedapproximately 21.2 elapsed accelerator-hours, excluding failures and local preprocessing.