Identifying similar words and contexts in natural language with SenseClusters

Ted Pedersen, Anagba Kulkarni · 2005

SenseClusters is a freely available intelligent system that clusters together similar contexts in natural language text. Thereafter it assigns identifying labels to these clusters based on their content. It is a purely unsupervised approach that is language independent, and uses no knowledge other than what is available in raw un-annotated corpora. In addition to clustering similar contexts, it can be used to identify syn-onyms and sets of related words. It has been applied to a di-verse range of problems, including proper name disambigua-tion, word sense discrimination, email organization, and doc-ument clustering. SenseClusters is a complete system that supports feature selection from large corpora, several differ-ent context representation schemes, various clustering algo-rithms, the creation of descriptive and discriminating labels for the discovered clusters, and evaluation relative to gold standard data.

Read the paper · More papers on PaperTik