Inference for probabilistic unsupervised text clustering

Loïs Rigouste, Olivier Cappé, François Yvon · IEEE/SP 13th Workshop on Statistical Signal Processing, 2005 · 2005

We investigate the use of a simple probabilistic model for unsupervised document clustering in large collections of texts. The model consists of a mixture of multinomial distributions over the word counts, each component corresponding to a different theme. The expectation-maximization (EM) algorithm is the basic tool used for inference. After introducing the model and experimental framework (corpus and evaluation measures), we discuss the importance of initialization and illustrate the difficulty caused by the lack of supervision information. We propose some ideas to solve this problem, one of the most efficient method being based on vocabulary reduction, and finally compare those heuristics with other inference processes, such as Gibbs sampling

Read the paper · More papers on PaperTik