Searching for topics in a large collection of texts

Martin Holub, Jiří Semecký, Jiří Diviš · 2004

We describe an original method that automatically finds specific topics in a large collection of texts. Each topic is first identified as a specific cluster of texts and then represented as a virtual concept, which is a weighted mixture of words. Our intention is to employ these virtual concepts in document indexing.In this paper we show some preliminary experimental results and discuss directions of future work.

Read the paper · More papers on PaperTik