A Methodology for Word Sense Disambiguation at 90% based on large-scale CrowdSourcing

Oier López de Lacalle, Eneko Agirre · 2015

Word Sense Disambiguation has been stuck for many years.In this paper we explore the use of large-scale crowdsourcing to cluster senses that are often confused by non-expert annotators.We show that we can increase performance at will: our in-domain experiment involving 45 highly polysemous nouns, verbs and adjective (9.8 senses on average), yields an average accuracy of 92.6 using a supervised classifier for an average polysemy of 6.1.Our proposal has the advantage of being cost-effective and being able to produce different levels of granularity.Our analysis shows that the error reduction with respect to finegrained senses is higher, and manual inspection show that the clusters are sensible when compared to those of OntoNotes and WordNet Supersenses.

Read the paper · More papers on PaperTik