Aggregating Continuous Word Embeddings for Information Retrieval

Stéphane Clinchant, Florent Perronnin · 2013

While words in documents are generally treated as discrete entities, they can be embedded in a Euclidean space which reflects an a priori notion of similarity between them. In such a case, a text document can be viewed as a bag-ofembedded-words (BoEW): a set of realvalued vectors. We propose a novel document representation based on such continuous word embeddings. It consists in non-linearly mapping the wordembeddings in a higher-dimensional space and in aggregating them into a documentlevel representation. We report retrieval and clustering experiments in the case where the word-embeddings are computed from standard topic models showing significant improvements with respect to the original topic models. 1

Read the paper · More papers on PaperTik