Text Processing Using Multilingual Resources at the Computing Research Laboratory
Jim Cowie · Literary and Linguistic Computing · 1994
The paper surveys five aspects of the work on multilingual text processing at the Computing Research Laboratory, New Mexico State University: a large-scale information extraction (IE) program; multilingual information retrieval; automatic document typing; a tagger for Spanish texts, and simulated annealing as a large-scale lexical disambiguator for text. These projects form what we hope is a coherent cluster of tools and applications directed towards the use of large-scale resources (corpora, machine-readable dictionaries, etc.) to build hybrid systems-statistical and symbolic computation hybrids—for a range of information extraction tasks. These range from IE proper (the extration of factual formatted information from texts) to improvements to conventional information retrieval tasks to more obviously linguistic ones, such as machine translation. The Computer Research Laboratory's emphases are also on the greatest degree of automation possible in the gathering and use of linguistic resources, and on the key role of active knowledge structures in the tasks.