Panel Discussion: "Evaluation Method of Machine Translation"

Muriel Vasconcellos · 1993

When it comes to credentials in MT evaluation, I have earned my stripes mainly as a frustrated observer of the process. I have watched MT evaluations from ALPAC to DARPA. As a hands-on user of MT for more than 13 years, I have seen and thought about the many forces and factors that come together to make MT effective. And I have learned how difficult they are to measure, espe-cially as they combine in countless different ways. It has always worried me to see hard-and-fast conclusions, sometimes sharply at odds with day-to-day experience, being drawn from isolated fragments of the picture, much as the apocryphal blind men felt different parts of the camel and made guesses about the whole animal that were widely off the mark. I am also a certified evaluee. During my watch, the MT project at the Pan American Health Organi-zation was subjected to six major studies. In the early years, three separate progress evaluations were done: by Wilks in 1978,by Zarechnak in 1981, and by Macdonald, also in 1981. In 1987, after seven years of practical operation, I was responsible for designing and implementing a controlled 11-month study that focused on cost-effectiveness and user satisfaction (Vasconcellos 1989). In 1989, the English-Spanish system, Engspan, was benchmarked against four others in a massive test conducted by McGraw-Hill (Benton 1989). And most recently, PAHO's Spanish-English sys-

Read the paper · More papers on PaperTik