Using Word Sense as a Latent Variable in LDA Can Improve Topic Modeling

Yunqing Xia, Guoyu Tang, Huan Zhao, Erik Cambria, Thomas Fang Zheng · 2014

Since proposed, LDA have been successfully used in modeling text documents. So far, words are the common features to induce latent topic, which are later used in document representation. Observation on documents indicates that the polysemous words can make the latent topics less discriminative, resulting in less accurate document representation. We thus argue that the semantically deterministic word senses can improve quality of the latent topics. In this work, we proposes a series of word sense aware LDA models which use word sense as an extra latent variable in topic induction. Preliminary experiments on benchmark datasets show that word sense can indeed improve topic modeling.

Read the paper · More papers on PaperTik