Document clustering and cluster topic extraction in multilingual corpora

Joaquim Silva, João Tiago Mexia, Carlos Agra Coelho, Gabriel Pereira Lopes · 2002

A statistics-based approach for clustering documents and for extracting cluster topics is described relevant (meaningful) expressions (REs) automatically extracted from corpora are used as clustering base features. These features are transformed and its number is strongly reduced in order to obtain a small set of document classification features. This is achieved on the basis of principal components analysis. Model-based clustering analysis finds the best number of clusters. Then, the most important REs are extracted from each cluster and taken as document cluster topics.

Read the paper · More papers on PaperTik