Sense Proximity versus Sense Relations

Julio A. Gonzalo · 2004

It has been widely assumed that sense distinctions in WordNet are often too ne-grained for applications such as Machine Translation, Information Retrieval, Text Classication, Document clustering, Question Answering, etc. This has led to a number of studies in sense clustering, i.e., collapsing sense distinctions in WordNet that can be ignored for most practical applications [1,5,6]. At the UNED NLP group, we have also conducted a few experiments in sense clustering with the goal of improving WordNet for Information Retrieval and related applications [4,3,2]. Our experiments led us to the conclusion that annotating WordNet with a typology of polysemy relations is more helpful than forming sense clusters based on a notion of sense proximity. The reason is that sense proximity depends on the application, and in many cases can be derived from the type of relation between two senses. In the case of metaphors, senses often belong to different semantic elds, and therefore a metaphor can be a relevant distinction for Information Retrieval or Question & Answer systems. For Machine Translation applications, however, the metaphoric sense extensions might be kept across languages, and therefore the distinction might not be necessary to achieve a proper translation. In the panel presentation, we will summarize the experiments that led us to hold this position: n In [2] we compared two clustering criteria: the rst criterion, meant for Information Retrieval applications, consists of grouping senses that tend to co-occur in Semcor documents. The second criterion, inspired by [7], groups senses that tend to receive the same translation in several target languages via the EuroWordNet Interlingual Index (parallel polysemy). The overlapping of both criteria was between 55% and 60%, which reveals a correlation between both criteria but leaves doubts about the usefulness of the clusters. However, a classication of the sense groupings according to the type of polysemy relation claries the data: all homonym and metaphor pairs satisfying the parallel polysemy criterion did not satisfy the co-occurrence criterion; all generalization/specialization pairs did satisfy the co-occurrence criterion; nally , metonymy pairs were evenly distributed between valid and invalid co-occurrence clusters. Further inspection revealed that the type of metonymic relation could be used to predict sense clusters for Information Retrieval. n In [3] we applied Resnik & Yarowsky measure to evaluate the Senseval-2 WordNet subset for sense granularity. We found that the average proximity was similar to the Senseval-1 sense inventory (Hector), questioning the idea that WordNet sense distinctions are ner than in other resources built by lexicographers. We also found that Resnik & Yarowsky proximity measure provides valuable information, but should

Read the paper · More papers on PaperTik