Comparative analysis of different word embedding models

Snehal Bhoir, Tushar H. Ghorpade, Vanita Manikrao Mane · 2017

The words in the vocabulary are mapped into real valued vectors of low-dimensional space which is relative to vocabulary size is termed as word embedding. Words with similar vectors are semantically similar. Words Embeddings are one of the few successful application of unsupervised learning with the major benefit that it do not require expensive annotations but they can be derived from large un-annotated corpora. In this paper, we present a comparative analysis of different word embedding models namely Continuous bag of words, Skip gram, Glove(Global Vectors for word representation) and Hellinger-PCA (Principal Component Analysis). The models are compared on different parameters. The parameters are performance with respect to size of training data, basic over-view, and relation of context and target words, memory consumption, supported classifier used and effect of changes in dimensionality. Word embedding turns text into numbers. This transformation has two important beneficial properties that is dimensionality reduction for efficient representation and contextual similarity for expressive representations.

Read the paper · More papers on PaperTik