Towards understanding word embeddings: Automatically explaining similarity of terms

Yating Zhang, Adam Jatowt, Katsumi Tanaka · 2016

Word embedding techniques (e.g., Word2Vec, GloVe) have been recently used for variety of applications with quite good rate of success. They allow to capture word semantics and syntactics with decreased dimensionality based on the concept of distributional vector representations. Vector representations can be then used for similarity comparison. However, if we treat the word embeddings as a kind of encryption process, it is difficult to decrypt their meaning. This makes it problematic to justify why particular terms should be considered similar as well as to prove that the overall quality of the trained vector space is high. Evaluating the accuracy of the similarity computation between any two given terms is difficult due to the lack of concrete evidences to explain and support the similarity. In this paper, we propose a novel way to automatically extract evidences represented as term pairs to explain the similarity of arbitrary terms. Our approach is unsupervised and can be applied to either homogeneous or heterogeneous vector spaces.

Read the paper · More papers on PaperTik