Semantic Similarity Analysis using FastText
Abhishek Rana, Akshita Pant, Nikita Rawat, Priyanshu Rawat, Satvik Vats, Vikrant Sharma · 2024
Accurately capturing semantic similarities between words is crucial for various natural language processing (NLP) applications, such as information retrieval and text summarization. Word embeddings, that represent words as dense vectors, have proven effective in this task. This study investigates the impact of dimensionality in association with data on the quality of FastText word embeddings and their ability to capture semantic similarities. Different FastText models (SkipGram and CBOW) are trained on a large Wikipedia corpus, with varying dimensionalities and are evaluated as per the Semantic Textual Similarity (STS) benchmark dataset using the Pearson Correlation Coefficient. Our findings provide insights into the model training time and performance, highlighting the strengths and limitations of different configurations, which helps in selecting appropriate dimensionality based on the desired balance between accuracy and computational efficiency. This research advances the understanding of the role of dimensionality with respect to data in word embedding models and their application to semantic similarity tasks, informing the development of robust NLP techniques for accurate language understanding.