A Simple Mixture Model For UnsupervisedText Categorisation

Fabrice Clérot, Françoise Fessant, Olivier Collin, Olivier Cappé, Éric Moulines · WIT transactions on information and communication technologies · 2004

Automatically segmenting text corpora into thematically related groups is a complex exploratory analysis problem. In this article, we outline our multi-stage exploratory analysis process and investigate the performance of a simple statistical model. After a description of this model and of its fitting procedure, we illustrate its performance on the segmentation of a corpus of CKM-related

Read the paper · More papers on PaperTik