Automatic WSD: Does it Make Sense of Estonian?
Kadri Vider, Kaarel Kaljurand · 2001
This paper describes a fully automatic Estonian word sense disambiguation system called semyhe which is based on Estonian WordNet (EstWN) hyponymjhypernym hierarchies and meant to disambiguate both nouns and verbs. 1 Short description of the system The main inspiration for our system is Agirre and Rigau (1996) similar system that disambiguates the English noun senses based on WordNet hyponymjhypernym hierarchy, taking into consideration the distances between the nodes corresponding to the word senses in the WordNet tree as well as the density of the tree. They have also experimented with using meronyms/holonyms in addition to hyponyms/hypernyms but report that it does not improve the results. Our main object was not to focus on the homonymous words only (lexical sample), but to try to disambiguate all nouns and verbs in the text. The Estonian WordNet (EstWN) also contains adjectives but they are not linked by hyponym/hypernym relations. The word sense disambiguation could also try to describe a unique sense for adverbs but in our case such words have not yet been included in the thesaurus. As far as we know this is the first attempt on automatic Estonian word sense disambiguation. 1.1 Input The input text for our system must be morphologically analyzed, meaning that each word is provided with its lemma and morphological reading. Taking those two into account we can localize the senses that correspond to the word in EstWN hyponym/hypernym tree (Vider et aL 1999). It must be mentioned that although the morphological description in the input can be quite detailed, we only use the information on whether the word is a noun or a verb. A simple morphological analysis that only looks at the word-form and not its context can result in very ambiguous output. On average 45 % of the words are morphologically ambiguous in Estonian texts (Kaalep, 1997). The ambiguity can be greatly reduced by also applying the Estonian morphological disambiguator (Kaalep and Vaino, 1998) to the text before the word sense disambiguation. Since even then the words can in principle stay morphologically ambiguous, our system doesn't require each word to have exactly one morphological reading assigned to it in the input text.