Subliminal Preference Transfer in LLM-Generated Training Data: A Theoretical Framework for Formalisation and Detection
Manasvi Gangrade · Zenodo (CERN European Organization for Nuclear Research) · 2026
This paper introduces Subliminal Preference Transfer (SPT), a theoretical framework describing how a teacher language model's latent behavioural preferences may become implicitly embedded within semantically neutral synthetic data, potentially influencing a student model trained on that data without explicit preference supervision. Unlike conventional data poisoning attacks that rely on anomalous content or trigger patterns, SPT is hypothesized to arise from subtle statistical regularities that remain within the distribution of otherwise benign data, making it difficult to detect using existing content-based filtering techniques. The paper formalizes SPT from a mechanistic interpretability perspective, models preference transfer as an alignment between latent preference representations and gradient-space dynamics during student training, and proposes SubtleNet, a gradient-space auditing framework for detecting potential indicators of SPT contamination in synthetic datasets. Rather than presenting empirical validation, this work establishes a conceptual framework, outlines an experimental methodology for future evaluation, and identifies open research questions surrounding synthetic data provenance, model alignment, and gradient-space safety auditing. A proof-of-concept prototype implementation is publicly available on GitHub. https://github.com/Manasvi-Gangrade/Subliminal-Preference-Transfer