Hungarian Word-Sense Disambiguated Corpus

Veronika Vincze, György Szarvas, Attila Almási, Dóra Szauter, Róbert Ormándi, Richárd Farkas, Csaba Hatvani, János Csirik · 2008

To create the first Hungarian WSD corpus, 39 suitable word form samples were selected for the purpose of word sense disambiguation. Among others, selection criteria required the given word form to be frequent in Hungarian language usage (frequency rates available in the Hungarian National Corpus (HNC) were used for measurement (Váradi, 2000)), and to have more than one sense considered frequent in usage. HNC and its Heti Világgazdaság (HVG) subcorpus provided the basis for corpus text selection. This way, each sample has a relevant context (the whole HVG article), and information on the lemma, POS-tagging and automatic tokenization is also available. 1. Word sense disambiguation Word Sense Disambiguation (WSD) aims at resolving ambiguities (homonymy, polysemy) in texts. This problem has been present in natural language processing (NLP) since the beginnings, and it is an important intermediate task for most NLP applications (e.g. text comprehension, human-machine interaction, machine translation and information retrieval and extraction). 1.1. Overview of previous research 1.1.1. Word sense disambiguation research in other languages Word sense disambiguation research concerning, first, English and, later, other languages as well was related in a greater part to SensEval (Kilgariff, 2001),(Mihalcea & Edmonds, 2004) workshops organized by ACL-SIGLex.

Read the paper · More papers on PaperTik