Multilingual corpus-based extraction and the Very Large Lexicon

Gregory Grefenstette · 2002

Over the past decade, the World Wide Web has been providing access to ever-increasing multilingual corpora. At the same time, computational linguists have been creating a wide range of linguistically motivated text abstraction techniques. These two phenomena permit the creation of extremely large collections of abstracted exemplars of text. One application of such exemplars is knowing the most likely abstracted form an extracted item would take, and another is to predict what the best translation for a term should be in a different language. In this article we describe the linguistic abstractions that a text can undergo, show how these abstractions can be stored in a Very Large Lexicon, and show one use of such a lexicon for multilingual term translation.

Read the paper · More papers on PaperTik