The Inverse Regression Topic Model

Maxim Rabinovich, David M. Blei · 2014

Taddy (2013) proposed multinomial inverse re-gression (MNIR) as a new model of annotated text based on the influence of metadata and re-sponse variables on the distribution of words in a document. While effective, MNIR has no way to exploit structure in the corpus to improve its predictions or facilitate exploratory data analy-sis. On the other hand, traditional probabilis-tic topic models (like latent Dirichlet allocation) capture natural heterogeneity in a collection but do not account for external variables. In this paper, we introduce the inverse regression topic model (IRTM), a mixed-membership extension of MNIR that combines the strengths of both methodologies. We present two inference algo-rithms for the IRTM: an efficient batch estima-tion algorithm and an online variant, which is suitable for large corpora. We apply these meth-ods to a corpus of 73K Congressional press re-leases and another of 150K Yelp reviews, demon-strating that the IRTM outperforms both MNIR and supervised topic models on the prediction task. Further, we give examples showing that the IRTM enables systematic discovery of in-topic lexical variation, which is not possible with pre-vious supervised topic models. 1.

Read the paper · More papers on PaperTik