Lessons Learned from a Task-based Evaluation of Speech-to-Speech Machine Translation

Lori S. Levin, Boris Bartlog, Ariadna Font Llitjós, DONNA M. GATES, Alon Lavie, Dorcas Wallace, Taro Watanabe, Monika Woszczyna · 2000

For several years we have been conducting Accuracy Based Evaluations (ABE) of the JANUS speech-to-speech MT system (Gates et al., 1997) which measure quality and fidelity of translation. Recently we have begun to design a Task Based Evaluation for JANUS (Thomas, 1999) which measures goal completion. This paper describes what we have learned by comparing the two types of evaluation. Both evaluations (ABE and TBE) were conducted on a common set of user studies in the semantic domain of travel planning. 1. Introduction For several years we have been conducting Accuracy Based Evaluations (ABE) (Gates et al., 1997) of the JANUS speech-to-speech machine translation system (Waibel, 1996; Levin et al., ). Our ABE focuses on whether the meaning of a source language segment is totally and accurately conveyed in the target language, and also includes a separate measure of fluency. This type of evaluation was useful in the early stages of system development for tracking our improvement over time...

Read the paper · More papers on PaperTik