The TALP systems for disambiguating WordNet glosses
Mauro Castillo, Francis J. Real, Jordi Asterias, Germán Rigau · Meeting of the Association for Computational Linguistics · 2004
This paper describes the TALP systems presented at Senseval-3 task 12 “Word-Sense Disambiguation of WordNet Glosses”. Our method combines a set of knowledge-based heuristics integrating several information sources and techniques. Using large scale lexico–semantic knowledge bases, such as WN, has become a usual, often necessary, practice for most current Natural Language Processing systems. Building appropriate resources of this nature for broad–coverage semantic processing is a hard and expensive task, involving large research groups during long periods of development. For example, dozens of person– years are been invested world–wide into the development of wordnets for various languages (Fellbaum, 1998), (Atserias et al., 1997), (Agirre et al., 2002), (Pianta et al., 2002). Dictionaries are special texts describing the meaning of a language. They provide a wide range of information of words by giving definitions of the word senses and as, a side effect, they supply knowledge about the world itself. WordNet (WN) (Fellbaum, 1998) can be also seen as an structured dictionary with thouthands of semantic relations, defining the most common concepts of the English language. Although the importance of (WN) has widely exceeded the purpose of its creation (Miller et al., 1990), and it has become an essential semantic resource for many applications, at the moment is not rich enough to directly support advanced semantic processing (Harabagiu et al., 1999). Sense disambiguation of definitions in any lexical resource is an important objective in the language engineering community because this process can increase the semantic conectivity among concepts. The first significant disambiguation of dictionary definitions took place 20 years ago (see (Rigau, 1998) for an extended survey on acquiring lexical knowledge from Machine Readable Dictionaries). Recently, several research groups have presented different approaches to perform this process on WN. In the eXtended WordNet1 (Mihalcea and Moldovan, 2001) the WN glosses have been syntactically parsed, transformed into logic forms and the content words are also semantically disambiguated. Being derived from an automatic process, disambiguated words included into the glosses have assigned a confidence label indicating the quality of the annotation (gold, silver or normal).