Word Network Topic Model Based on Word2Vector
Mingyi Jiang, Rui Liu, Fei Wang · 2018
With the rapid development of social media and e-commerce on the Internet, many short texts such as instant messages, tweets and product comments have become an important form of Internet information. Compared with regular texts, short texts suffer severe data sparsity and informal format. Because of that, mining latent semantics such as topics in short texts is a critical challenge. There are many limitations directly applying common topic models like Latent Dirichlet Allocation (LDA) to short texts. In result of that, many short text topic models were put forward. The Word Network Topic Model (WNTM) is the typical one among them, which transfers sparse document to word space into dense word to word space and then learn topics from it. However, WNTM has its own limitation. The word co-occurrence matrix of the original document is used to construct the word network. This method is simple and intuitive, which can't express the deep meaning between words and words. To solve this problem, we proposed W2V-WNTM, which uses Word2Vec model instead of word co-occurrence to construct word network. Experiment shows it can describe relationship between words better.