When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Edits
Jabez Magomere, Emanuele La Malfa, Manuel Tonneau, Ashkan Kazemi, Scott A. Hale · 2025
Online misinformation remains a critical challenge, and fact-checkers increasingly rely on claim matching systems that use sentence embedding models to retrieve relevant fact-checks.However, as users interact with claims online, they often introduce edits, and it remains unclear whether current embedding models used in retrieval are robust to such edits.To investigate this, we introduce a perturbation framework that generates valid and natural claim variations, enabling us to assess the robustness of a wide-range of sentence embedding models in a multi-stage retrieval pipeline and evaluate the effectiveness of various mitigation approaches.Our evaluation reveals that standard embedding models exhibit notable performance drops on edited claims, while LLM-distilled embedding models offer improved robustness at a higher computational cost.Although a strong reranker helps to reduce the performance drop, it cannot fully compensate for first-stage retrieval gaps.To address these retrieval gaps, we evaluate train-and inference-time mitigation approaches, demonstrating that they can improve in-domain robustness by up to 17 percentage points and boost out-of-domain generalization by 10 percentage points.Overall, our findings provide practical improvements to claim-matching systems, enabling more reliable fact-checking of evolving misinformation.Mitigation Approaches Covid-19 no pass ordinary flu for how e dey kill people Covid-19 no pass ordinary flu for how e dey kill people Covid-19 is no more deadly than the ordinary flu.Teacher Model Student Model Covid-19 is only as deadly as the seasonal flu MSE-Loss MSE-Loss Knowledge Distillation Claim Normalization q q' q' q'' Input Claim (q): Covid-19 is only as deadly as the seasonal flu Fact Check: COVID-19 has a significantly higher mortality rate than seasonal flu, with greater severity and hospitalizations LLM As a Perturber LLM As a Verifier Perturbation Generation N candidate rewrites verified perturbations COVID-19 IS ONLY AS DEADLY AS THE SEASONAL FLU Covid-19 no pass ordinary flu for how e dey kill people COVID-19 is onli as dedli as the siznl flu Casing Typos Negation Entity Replacement LLM Rewrite Dialect Stage Model TrueCase Upper Least Most Shallow Double Atleast 1 All Least Most AAE Jamaican Pidgin Singlish CheckThat22 First-Stage Retrieval BM25 +0.0 +0.0 -15.2 -15.0 +0.5 -2.6 -2.7 -12.4 -0.8 -0.2 -9.2 -6.7 -5.0 -1.6 all-distilroberta-v1 -0.4 -35.4 -15.8 -13.9 +1.8 +2.4 -3.0 -5.4 +3.5 +5.3 -4.4 -11.4 -6.6 -0.7 all-MiniLM-L12-v2 +0.0 +0.0 -8.2 -13.6 +3.5 +4.4 -2.9 -2.5 +3.1 +4.4 -6.7 -10.0 -5.8 +0.2 all-mpnet-base-v2 +0.0 +0.0 -8.3 -8.4 +0.2 -4.2 -3.1 -9.8 +1.0 +1.8 -3.8 -7.5 -8.5 -1.3 all-mpnet-base-v2-ft +0.0 +0.0 -8.6 -12.8 -1.4 -4.2 -1.8 -8.6 -1.3 -1.7 -3.8 -4.0 -8.5 -3.2 sentence-t5-base -2.5 -15.8 -2.9 -7.1 -21.4 -12.2 -3.3 -5.7 +2.3 +2.0 -0.9 -8.2 -5.5 -0.1 sentence-t5-large -1.0 -9.9 -2.1 -5.4 -27.2 -14.7 -0.9 -5.4 +1.8 +1.7 +0.8 -3.4 -3.2 +2.0 sentence-t5-large-ft -0.6 -9.7 -1.8 -3.6 -11.7 -6.6 -2.5 -1.6 +2.2 +1.0 +1.5 -2.4 -3.7 +1.2 instructor-base -1.1 -9.4 -4.4 -3.2 -3.7 -3.8 -1.7 -8.0 +0.1 -0.8 -1.1 -4.0 -4.3 -1.5 instructor-large -0.7 -4.5 -1.1 -1.7 -4.1 -4.9 +0.1 -2.5 +0.0 -0.9 -0.9 -2.5 -4.1 -2.0 SFR-Embedding-Mistral -0.6 -0.6 -0.7 -0.5 -1.2 -2.9 -0.8 -1.1 +0.1 +0.7 -1.9 -0.9 -1.2 -2.2 NV-Embed-v2 +0.0 -0.6 +0.7 +-0.0 -0.5 -2.4 +0.2 -0.3 -0.1 -0.2 -1.0 -0.2 -0.9 -1.2