Unlinkable but Not Anonymous: A Low-Cost Front-End that Breaks Cross-Session Linkage of a Deterministic Decision Token

Randolph James Ferlic, Kimberly Kate Ferlic · Zenodo (CERN European Organization for Nuclear Research) · 2026

Unlinkable but Not Anonymous: A Low-Cost Front-End that Breaks Cross-Session Linkage of a Deterministic Decision Token Randolph James Ferlic, M.D. and Kimberly Kate Ferlic (Fieldstone Analytics, LLC, Austin, TX, USA) Preprint · Zenodo DOI: 10.5281/zenodo.22838120 · CC-BY 4.0 · Community: spiral-domain-encoder-campaign Abstract A companion characterization established that a compact, deterministic edge "decision token" — a window of a sensor stream quantized on-device to a single ~8-bit code — is non-invertible for the exact signal but not anonymizing: it re-identifies the individual, links that individual across recordings (cross-session 0.71 on ECG), and leaks demographic attributes. It is a pseudonym, not an anonym. This paper asks the operational follow-on: can a fixed, cheap, filed front-end — a decision-orthogonal identity-axis projection that removes an estimated per-individual subspace before quantization — reduce that identifiability, and does the reduction survive a strong attacker? We report a sharp, honest split, measured against a battery of attackers (logistic, kNN, SVM, random forest, gradient boosting, MLP). (i) It does NOT confer within-session anonymity. Against a linear attacker the projection appears to de-identify (re-id 0.99 → 0.40), but this is a weak-attacker artifact — a nonlinear attacker recovers the individual almost entirely (~0.98), because linear subspace removal leaves the identity in the nonlinear residual. (ii) But it DOES confer cross-session unlinkability. The durable, deployment-relevant threat — linking an individual across recordings — is robustly broken: the strongest attacker's cross-session re-identification falls from 0.71 to 0.19 (n = 26 patients) and from 0.22 to 0.09 (a 61% reduction of the above-chance value) at n = 220, holding against nonlinear attackers and as the gallery grows. The front-end removes the session-transferable identity, which is exactly what cross-session linkage needs. (iii) It reduces demographic attribute inference only partially (sex 0.625 → 0.560, age 0.484 → 0.419) and (iv) its decision cost is task-dependent (zero on an identity-orthogonal diagnosis, ~0.05 AUROC on a harder one). Separately, the linearly-removable identity is separable from the token's distributional cross-individual drift, which the projection does not fix. The precise, standards-aligned conclusion (ISO/IEC 24745): a fixed linear front-end delivers measured unlinkability, not anonymity — the token remains identifying data whose cross-session linkability is cheaply and robustly reducible. This is a characterization of a previously-filed method and discloses no new algorithmic subject matter. Highlights · The honest, strong-attacker result. Every re-identification claim is tested against six attackers, not one — because a privacy defense is only as strong as the best attack against it. This catches the weak-attacker failure mode that a linear-only evaluation would miss. · Within-session anonymity FAILS. The tempting headline "de-identification 0.99 → 0.40" is a linear-attacker artifact: a nonlinear attacker (random forest / gradient boosting / MLP) recovers the individual at ~0.98 from the identical projected features. Linear subspace removal does not confer within-session anonymity. · Cross-session UNLINKABILITY holds — robustly. The front-end breaks the durable threat — linking a person across recordings — even against strong attackers and as the gallery grows: strongest-attacker cross-session re-id 0.71 → 0.19 (n = 26) and 0.22 → 0.09 (−61% above chance, n = 220). It removes the session-transferable identity, which is what cross-session linkage depends on. · Standards-aligned framing (ISO/IEC 24745). The result is measured unlinkability, not irreversibility-grade anonymity — the constructive complement to the companion "Non-Invertible but Not Anonymous." The right claim is confidentiality + minimization strengthened by unlinkability, never anonymization. · Honest bounds. Attribute inference is reduced only partially (not an attribute scrubber); the decision cost is task-dependent (zero on identity-orthogonal decisions, ~0.05 on identity-adjacent ones), so the number of removed directions is a privacy ↔ accuracy knob. · A clean scientific point. The linearly-removable identity is separable from the token's distributional cross-individual drift — the front-end is a linkability tool, not a generalization tool. · Population scale. Pushed to 2,094 multi-recording patients and 6,759 within-session identities: the 1-byte entropy cap is the primary unlinkability driver (single-token re-identification collapses toward — though stays measurably above — chance as the gallery grows, 0.44 → 0.011 within-session, 0.16 → 0.02 cross-session), with the front-end a complementary ~18% reduction on top. The unlinkability claim rests on a property of the token itself, not only the front-end. · The stream, if long enough, defeats the cap. The single-token entropy cap protects a snapshot (0.170→0.003 to the full ~21.8k population) but the token STREAM accumulates identity ~linearly with length: a 20-window stream re-identifies 0.121 at n=10k (crossing the 0.10 line), so a long monitoring stream re-identifies at population scale — reinforcing keep-the-token-ephemeral / egress-only-the-decision. · Cross-session leakage generalizes across modalities. Same protocol on EEG (109 subjects, cross-run), glucose (112 patients, cross-day), sEMG, robotics, industrial: identity leakage is universal and CI-significant (P>0.98); EEG is the strongest non-ECG biometric (0.578) and the front-end contains it where n permits (0.578→0.486); honest bound — none exceed ECG's n, a generalization not a bigger-n result. · Co-channels multiply identity leakage. Emitting multiple token co-channels (per-lead / per-channel tokens — the filed channel-partition ladder) is far more identifying than one: ECG K=6 re-identifies from just 2 windows (0.527 at n=10k) where a single channel can't; universal across wearables — a 64-channel EEG headset's per-channel tokens reach 0.890 (97× chance). The emitted-channel count is itself a privacy budget: compute co-channels on-device, egress only the single decision. · Pre-registered, and reported against the strongest attacker. Every claim here is the one that survives a six-model attacker battery: the weak-attacker "de-identification" reading does not, and the durable cross-session unlinkability does. What this record contains · Manuscript_Paper45.pdf — the manuscript with seven figures embedded (anonymity-fails / unlinkability-holds; separability; attribute + cost; population-scale entropy cap; stream-length; multi-modal generalization; co-channel emission), and Manuscript_Paper45.docx, the editable source. · PAPER_45_ZENODO_ARCHIVE.zip — the reproducibility archive: the frozen pre-registrations (`IDSUPPRESS`, `DESUPPRESS`, `HARDEN_PREDEPOSIT`, `HARDEN_STREAM`, `POPSCALE`, `STREAMSCALE`, `XMODAL`, `COCHAN`), the runners (identity-suppression sweep, PTB-XL cross-session/attribute/cost, and the strong-attacker battery within-session and cross-session at n = 26 and n = 220, plus the stream/feature/k-sweep and attribute-battery stress tests, and the population-scale entropy-cap / cross-session-at-scale runners), the per-experiment result records, the seven figures, the manuscript source, and a README. All datasets are public; no raw benchmark data is redistributed (sources and a `PATH_TO_DATA` convention are in the README). All paths/identifiers are scrubbed and leak-scanned per the campaign deposit discipline. Cite as R. J. Ferlic and K. K. Ferlic, "Unlinkable but not anonymous: a low-cost front-end that breaks cross-session linkage of a deterministic decision token," Zenodo, 2026, doi: 10.5281/zenodo.22838120. License and patent notice Released under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). Consistent with that license, no patent, patent application, or other intellectual-property right of the authors is licensed, waived, granted, or otherwise conveyed by this deposit. This work characterizes a previously-filed method and discloses no new algorithmic subject matter. The methods characterized — the class-discriminant single-token codebook encoder and its nearest-centroid monitor, and the decision-orthogonal identity-axis-projection front-end that is the subject of this paper — are the subject of filed and pending U.S. patent applications held by the authors, including U.S. Provisional Application No. 64/095,354 (the encoder), U.S. Provisional Application No. 64/084,821 (identity-axis projection, the front-end characterized here), and U.S. Non-Provisional Application No. 19/691,234 (privacy-preserving operations on encoded states). Per-deployment productization and deployment-selection know-how are not disclosed and are retained as trade secrets. © 2026 Fieldstone Analytics, LLC and the authors; all rights not expressly granted under CC-BY 4.0 are reserved. Licensing and collaboration inquiries: [email protected]. Companion deposits (spiral-domain-encoder-campaign) · Non-invertible but not anonymous: a privacy characterization of the token (the negative this paper answers): doi:10.5281/zenodo.22819210 · The price of the bottleneck: a multi-domain deployment characterization of the token (co-deposited companion): doi:10.5281/zenodo.22838118 · Class-discriminant single-token codebook construction (the encoder): doi:10.5281/zenodo.20788187 · Deterministic multi-token token ladder: doi:10.5281/zenodo.22003179 · Label-free inference-time channel fusion: doi:10.5281/zenodo.22046713 · A decision-oriented token as a bounded, threshold-free cache key for generative edge outputs: doi:10.5281/zenodo.22148612 · The predictive reach of a decision token: doi:10.5281/zenodo.22736921 Keywords edge AI priv

Read the paper · More papers on PaperTik