DutchSemCor: in quest of the ideal sense-tagged corpus

Piek T. J. M. Vossen, Rubén Izquierdo, Attila Görög · VU Research Portal · 2013

The most-frequent-sense and the predominant domain sense play an important role in the debate on word-sensedisambiguation.This discussion is, however, biased by the way sense-tagged corpora are built.In this paper, we argue that current sense-tagged corpora neglect rare senses and contexts and, as a result, do not represent a good corpus for training and testing word-sensedisambiguation.We defined three quality criteria for sense-tagged corpora and a methodology to satisfy these criteria with minimal effort.Following this method, we built a Dutch sense-tagged corpus that tried to meet these criteria.The corpus was evaluated by deriving word-sensedisambiguation systems and testing these on different subsets of the corpus in different ways.The performance of our systems and the quality of the derived data are equal to state-of-the-art English systems and corpora.Finally, we used the systems to create a Dutch corpus of over 47 million sense-tagged tokens spread over a large variety of genres, domains and usages of Dutch.The results of the project can be downloaded freely from the project website.

Read the paper · More papers on PaperTik