Addressing Asymmetry in Contrastive Learning: LLM-Driven Sentence Embeddings with Ranking and Label Smoothing
Yan Huang, Shaoben Zhu, Wei Liu, Jiayi Wang, Xinheng Wei · Symmetry · 2025
Unsupervised sentence embedding, vital for numerous NLP tasks, struggles with the inherent asymmetry of semantic relationships within contrastive learning (CL). This paper proposes Label Smoothing-based Ranking Negative Sampling (LS-RNS), a novel framework that directly tackles the semantic asymmetry between anchor and negative samples in CL. LS-RNS utilizes a Large Language Model (LLM) to assess fine-grained asymmetric similarity scores between sentences, constructing a ranking-aware negative sampling strategy combined with adaptive label smoothing. This design encourages the model to learn more effectively from informative negatives that are semantically closer to the anchor, leading to asymmetry-aware sentence embeddings. Experiments on standard Semantic Textual Similarity (STS) benchmarks (STS12–STS16, STS-B, SICK-R) show that LS-RNS achieves state-of-the-art performance. We adopt Spearman’s rank correlation coefficient as the primary evaluation metric for semantic similarity tasks, and we use classification accuracy for downstream and transfer tasks. LS-RNS achieves 79.87 on STS tasks with BERT-base (vs. 76.25 for SimCSE, +3.62) and 80.41 with RoBERTa-base (vs. 79.18 for DiffCSE). On transfer tasks, it attains 88.82 (BERT) and 87.68 (RoBERTa), consistently outperforming PromptBERT and SNCSE. On STL-10, LS-RNS improves SimCLR top-one accuracy from 79.50% to 80.52% with ResNet-18 and from 68.91% to 72.19% with VGG-16, even enabling a shallow ResNet-18 to surpass a deeper ResNet-34 baseline. These results confirm the modality-agnostic effectiveness of LS-RNS and its potential to redefine contrastive learning objectives by modeling semantic asymmetry, rather than relying solely on encoder depth or pre-training objectives.