Text Semantic Analysis Algorithm Based on LDA Model and Doc2vec
Wei Wei Zhang, Guangyu Zhai, Binbin Zhong, Xiaoyi Kong · Advances in transdisciplinary engineering · 2024
With the rapid development of the Internet, massive messy text data is distributed in all walks of life, how to quickly and effectively mine meaningful semantic information from these complex and messy texts has become an important task in the field of natural language. First, texts such as Weibo comments lack contextual semantics due to sparse data, which in turn affects the effect of text semantic mining. Second, existing topic model use semantic representations in different vector spaces, resulting in low accuracy. Therefore, this paper proposes a new text semantic analysis model: DBOW-LDA model, which integrates the LDA topic model (Latent Dirichlet Allocation) and the sentence vector model (Doc2vec), so that the output contains the given topic semantic information. Sentence vector representation of text. The accuracy of the algorithm is further improved. In the experiment, the crawled Weibo comment text is used as the data set, and the K-Means clustering algorithm is used to compare the effects of each model. The experimental results show that the clustering effect based on DBOW-LDA model is better than Word2vec, WT-LDA, LDA+Word2vec model.