Sense Clustering Using Wikipedia

Bharath Dandala, Chris Hokamp, Rada F. Mihalcea, Răzvan Bunescu · Recent Advances in Natural Language Processing · 2013

In this paper, we propose a novel method for generating a coarse-grained sense inventory from Wikipedia using a machine learning framework. Structural and content-based features are employed to induce clusters of articles representative of a word sense. Additionally, multilingual features are shown to improve the clustering accuracy, especially for languages that are less comprehensive than English. We show the effectiveness of our clustering methodology by testing it against both manually and automatically annotated datasets.

Read the paper · More papers on PaperTik