Beyond semantics: Content leakage mitigation using synthetic hard negatives for style embeddings

Javier Huertas‐Tato, Adrián Girón-Jiménez, Alejandro Martín, David Camacho · Expert Systems with Applications · 2026

Purpose: Authorship can be defined as a combination of content and style . Modern open-source transformer foundational authorship models apply contrastive learning techniques. When naively contrasting texts to an authorship task some amount of semantic leakage is present, as authors frequently repeat their topic preferences. Our aim is to reduce spurious correlations due to topic leakage born from contrastive objectives. Methodology: We present a technique to modify a well-established contrastive learning objective (InfoNCE) using synthetic hard negative examples for in-domain topic leakage improvements, while preserving competitive out of domain performance. This topic leakage mitigation technique aims to distance the content embedding space from the style embedding space. Our experiments aim to demonstrate this technique in detail, using exclusively affordable encoder-only models instead of costly hard negative mining. Results: We showcase the performance with ablations on two different datasets and compare them on out-of-domain challenges. We improve on challenging evaluations with prolific authors, with up to 10% increase in accuracy on highly diverse bodies of work . Trials with standard challenges also demonstrate the preservation of zero-shot capabilities of this method as fine-tuning.

Read the paper · More papers on PaperTik