Cosine Similarity-Based Evidences Selection for Fact Verification Using SBERT on the FEVER Dataset

Harya Gusdevi, Arief Setyanto, Kusrini Kusrini, Ema Utami · CogITo Smart Journal · 2025

The spread of misinformation on digital platforms has emphasized the urgent need for automated fact verification systems. However, selecting the most semantically relevant evidence to support or refute a claim remains a challenge, especially within the widely used FEVER dataset. Traditional approaches like TF-IDF often fall short in capturing the contextual meaning between claims and evidence. This study addresses the problem by comparing TF-IDF with Sentence-BERT (SBERT) in measuring semantic similarity. The novelty of this research lies in embedding both claims and evidence using SBERT, then calculating cosine similarity to quantify their semantic relevance. Before embedding, standard preprocessing steps were applied, including tokenization, stemming, lowercasing, and stopword removal. A quantitative approach is used to compute cosine similarity between claim-evidence pairs using both TF-IDF and SBERT embeddings. Similarity analysis, distribution statistics, and t-tests are conducted to evaluate the methods. The results show that SBERT achieves higher similarity with the “SUPPORTS” category (0.65) and stronger negative similarity with “NOT ENOUGH INFO” (-0.90), compared to TF-IDF (0.49 and -0.62, respectively). SBERT also demonstrates more stable score distributions and significantly higher t-test values across all label comparisons, indicating stronger semantic discrimination. These findings confirm that SBERT outperforms TF-IDF in identifying the most relevant evidence. The new dataset generated can serve as a foundation for future fact verification model development.

Read the paper · More papers on PaperTik