Automatic Acquisition of Sense Tagged Corpora

Rada F. Mihalcea, DAN I. MOLDOVAN · 1999

An important problem in Natural Language Processing is identifying the correct sense of a word in a particular context. Thus far, statistical methods have been considered the best techniques in word sense disambiguation. Unfortunately, these methods produce high accuracy results only for a small number of preselected words. The reduced applicability of statistical methods is due basically to the lack of widely available semantically tagged corpora. In this paper we present a method which enables the automatic acquisition of sense tagged corpora. It is based on (1) the information provided in WordNet, particularly the word denitions found within the glosses and (2) the information gathered from Internet using existing search engines. Introduction Word Sense Disambiguation (WSD) is an open problem in Natural Language Processing. Its solution impacts other tasks such as discourse, reference resolution, coherence, inference and others. Thus far, statistical methods have been considered ...

Read the paper · More papers on PaperTik