DutchSemCor: Targeting the ideal sense-tagged corpus

Piek T. J. M. Vossen, Attila Görög, Rubén Izquierdo, Antal van den Bosch · 2012

Word Sense Disambiguation (WSD) systems require large sense-tagged corpora along with lexical databases to reach satisfactory results.The number of English language resources for developed WSD increased in the past years while most other languages are still under-resourced.The situation is no different for Dutch.In order to overcome this data bottleneck, the DutchSemCor project will deliver a Dutch corpus that is sense-tagged with senses from the Cornetto lexical database.In this paper, we discuss the different conflicting requirements for a sense-tagged corpus and our strategies to fulfill them.We report on a first series of experiments to support our semi-automatic approach to build the corpus.

Read the paper · More papers on PaperTik