Investigating topic coherence and task performance for varied types of words in LDA

Chuen‐Min Huang · 2017

In this study, we compared three types of word assignment to investigate topic coherence and task performance in Latent Dirichlet Allocation (LDA). We randomly selected 18,000 news articles consists of the categories of finance, life, and politics with 176,224 unigram, 863,680 compound, and 938,206 mixture in total. The result shows that our unigram-based model has high interpretability and topic coherence which the traditional LDA has been blamed for its shortage. The compound and mixture models illustrate more precise and accurate meaning and demonstrate computational efficiency than the unigram based model. This result suggests a well-planned word preprocessing is a crucial factor for supporting the superior topic clustering function of LDA.

Read the paper · More papers on PaperTik