Enhancing Short Text Semantic Similarity Measurement Using Pretrained Word Embeddings and Big Data
Supakpong Jinarat, Ratchakoon Pruengkarn · 2024
Measuring semantic similarity between short texts is a fundamental task in natural language processing with applications in information retrieval, question answering, and machine translation. Traditional methods such as term frequency (tf) and term frequency-inverse document frequency (tfidf) rely on lexical matching and fail to capture semantic meanings. This paper introduces Word Embeddings for Semantic Similarity (WESS), leveraging pretrained Word2Vec embeddings to capture semantic relationships. We fine-tune the Word2Vec model on the Quora Question Pairs dataset, containing 404,290 pairs of questions labeled as duplicate or non-duplicate. Our approach calculates text similarity by aggregating word embedding similarities. Experimental results demonstrate that WESS outperforms traditional methods, achieving an accuracy of 0.675, a 9.4% improvement over tf (0.617), a 6.5% improvement over tfidf (0.634), a 3.8% improvement over tfidf combined with Word2Vec (0.636), and a 1.9% improvement over standalone Word2Vec (0.662). These findings underscore the importance of semantic understanding in text similarity tasks and validate the effectiveness of pretrained word embeddings for capturing nuanced semantic relationships. The SSST method offers a robust and accurate approach for measuring semantic similarity between short texts, providing significant improvements over traditional approaches.