A Simple Mixture Model For UnsupervisedText Categorisation
Fabrice Clérot, Françoise Fessant, Olivier Collin, Olivier Cappé, Éric Moulines · WIT transactions on information and communication technologies · 2004
Automatically segmenting text corpora into thematically related groups is a complex exploratory analysis problem. In this article, we outline our multi-stage exploratory analysis process and investigate the performance of a simple statistical model. After a description of this model and of its fitting procedure, we illustrate its performance on the segmentation of a corpus of CKM-related