A Management Tool for Test Corpora

Gerardo Arrarte, Teófilo Redondo, Miguel Sobejano, Isabel María Castro Zapata · 1991

IBM is engaged in advanced research and development projects on various aspects of Machine Translation between several language pairs, one of which, LMT (Logic-based Machine Translation), was started by M. C. McCord in T. J. Watson Research Center in 1985 (McCord 85, 88, 89a, 89b, 89c). The IBM Madrid Scientific Center is taking part in this project through the development of English-Spanish and Spanish-English MT systems. One of the major challenges in all MT systems is testing and validating the system performance through the use of linguistic corpora (King 90). For the English-Spanish LMT prototype a large and varied set of corpus sentences (around 7500, at this stage) has been selected for the purpose of testing. The sentences were taken from different IBM hardware and software manuals, as well as some other sources such as dictionaries, grammar books and others. Finally, made-up sentences have also been included to cope with specific English-Spanish translation problems. We have made a special point of developing a tool for building and managing such corpora: the Corpus Database Manager (CDBM). CDBM enables NLP researchers to have ready access to the texts in the corpora in a selective way. That means being able to select sentences sharing linguistic features which are implicit in them, but which need to be stated by a linguist. To do this, sentences had to be marked first with a set of labels showing each of the features they contain. Both the selection (or retrieving) and classification (or typification) processes imply a huge amount of work by qualified experts, which could be optimized by means of the CDBM. This tool takes advantage of Database facilities to store the sentences along with the labels attached to them.

Read the paper · More papers on PaperTik