Neural Topic Models for Short Text Using Pretrained Word Embeddings and Its Application To Real Data

Riki Murakami, Basabi Chakraborty · 2021

Latent Dirichlet Allocation (LDA) is a typical example of a topic model that estimates the latent topics of sentences. It is widely used in topic discovery, information retrieval, and document modeling. In recent years, with the advancement of research about neural networks, topic models using neural networks, such as NVLDA and ProdLDA, have been presented. Since it is easy to use accelerators such as GPUs for training these, topic modeling for large corpora can be done efficiently. Topic models are also used for short texts such as social networking sites and product reviews, where the number of words in a document is often short. This means that the co-occurrence information of words is also extremely less, and it is difficult to infer latent topics for easy understanding by humans when training with only the target corpus. In this paper, we investigated whether trained word embedding vectors on other large corpora such as Wikipedia compensate for the lack of word information in short texts. While previous studies have been conducted on long texts such as newspaper articles, we specifically checked the effect on short texts. As a result, we confirmed that under many conditions, Topic Coherence evaluation using Wikipedia was improved by using word embedding. However, we were not able to achieve stable high performance in terms of topic diversity. We also used this approach for topic modeling of tweets related to COVID-19.

Read the paper · More papers on PaperTik