Combining statistics and semantics for word and document clustering

Alexandre Termier, Marie-Christine Rousset, Michèle Sébag · 2001

A new approach for constructing pseudo-keywords, referred to as Sense Units, is proposed. Sense Units are obtained by a word clustering process, where the underlying similarity reflects both statistical and semantic properties, respectively detected through Latent Semantic Analysis and WordNet.

Read the paper · More papers on PaperTik