Mongolian Information Retrieval Method Based on Word2vec and Topic Model

Siriguleng · 2019

To capture the real intention of users' needs more accurately from the increasingly abundant Mongolian information and return the retrieval results that best meet their needs, a Mongolian information retrieval method based on Word2vec and LDA topic model is proposed in this paper. Combining Mongolian grammatical features, this method builds a model based on LDA three-tier Bayesian structure to mine the hidden topic distribution and feature word distribution of documents, expands user queries according to Word2vec model to obtain words similar to user query keywords semantically, and then uses topic model to model extended vocabulary. Finally, according to the user's query topic, the similarity between the query topic and the document topic is calculated, and the document with high relevance to the query topic is returned. The experimental results show that the effective combination of Word2vec and LDA model achieves better results than the traditional model with initial query in the representation of latent semantics.

Read the paper · More papers on PaperTik