Bag of Embedded Words learning for text retrieval
Nikolaos Passalis, Anastasios Tefas · 2016
The word embedding models are capable of capturing the semantic content of the textual words. The process of extracting a set of word embedding vectors from a text document is similar to the feature extraction step of the Bag-of-Features pipeline, which is usually used in computer vision tasks. That gives rise to the Bag-of-Embedded Words (BoEW) model. In this paper a novel learning technique that optimizes both the word embedding and the codebook of the BoEW model towards text retrieval is proposed. The proposed method adheres to the cluster hypothesis that states that points in the same cluster are likely to fulfill the same information need and it is demonstrated, using two text datasets, that can significantly increase the retrieval precision. Finally, the proposed technique uses smaller representations than the competitive representation methods, that allows to reduce both the retrieval time and the storage requirements.