Making Topic Models more Usable

Wray Buntine · 2015

The output of topic models has always been seductive but not quite satisfying ever since the early work of Hofmann (PLSI) and Lee and Seung (NMF). An important approach to cleaning up the semantics does output analysis using coherence, and indeed other document summarization methods could also be used. However, this talk argues that topic models themselves need attention. New ways of modelling document semantics are being explored in the field of deep neural networks. Similarly, non-parametric versions of topic models allow modelling such effects as document structure, word sparsity, word burstiness, background words, multi-word terms, and network effects from author or follower networks, and semantic word hierarchies. These are usually done in the spirit of deep neural networks using hierarchical models, but earlier algorithms were often too slow to be realistic.

Read the paper · More papers on PaperTik