A Text Retrieval Algorithm Based on the Hybrid LDA and Word2Vec Model

Mu Xue · 2019

Text retrieval is a fundamental problem in information retrieval, and it is greatly influenced by the text feature representation. In this paper, we propose a novel text feature representation model, which integrates the LDA and Word2Vec model, and then utilize this model to handle the text retrieval problem. Latent Dirichlet Allocation represents a generative probabilistic model of a corpus, and texts are organized as random mixtures of latent topics. In particular, the proposed algorithm computes the distance between document and topics, and then each document is represented as a feature vector, in which each dimension denotes the distance between this document and a specific topic. To test the effectiveness of the proposed algorithm, several related methods are made performance comparison, and experimental results demonstrate that the proposed solution performs better than other methods, and it can achieve high accuracy for text retrieval.

Read the paper · More papers on PaperTik