Critical Dimension of Word2Vec
Shuvayanti Das, Sohan Ghosh, Shubham Bhattacharya, Rajat Varma, Dinabandhu Bhandari · 2019
Word embeddings are an efficient way of representing text such that they can be used by different Machine Learning Algorithms. Word2Vec is one such word embedding model. Although it is highly efficient, this model can take up a lot of space to store the vector representations. So defining the appropriate dimension of this model is very important to improve its performance in memory restricted devices. In this work, we present an empirical approach to decide the dimension of the word embeddings for a specific set of documents (corpora) to a critical value such that the representation of the words still preserve their original semantic and syntactic meanings. In this context, principal component analysis (PCA) has been used to compare the precision of vectors represented in different dimensions.