Dirichlet mixtures in text modeling

Mikio Yamamoto, Kugatsu Sadamitsu · Institutional Repositories DataBase (IRDB) · 2005

Word rates in text vary according to global factors such as genre, topic, author, and expected readership (Church and Gale 1995).Models that summarize such global factors in text or at the document level, are called 'text models.'A finite mixture of Dirichlet distribution (Dirichlet Mixture or DM for short) was investigated as a new text model.When parameters of a multinomial are drawn from a DM, the compound for discrete outcomes is a finite mixture of the Dirichlet-multinomial.A Dirichlet multinomial can be regarded as a multivariate version of the Poisson mixture, a reliable univariate model for global factors (Church and Gale 1995).In the present paper, the DM and its compounds are introduced, with parameter estimation methods derived from Minka's fixed-point methods (Minka 2003) and the EM algorithm.The method can estimate a considerable number of parameters of a large DM, i.e., a few hundred thousand parameters.After discussion of the relationships within the DM -probabilistic latent semantic analysis (PLSA) (Hofmann 1999), the mixture of unigrams (Nigam et al. 2000), and latent Dirichlet allocation (LDA) (Blei et al. 2001(Blei et al. , 2003) ) -the products of statistical language modeling applications are discussed and their performance in perplexity compared.The DM model achieves the lowest perplexity level despite its unitopic nature.

Read the paper · More papers on PaperTik