A proposal for an evaluation set selection method for corpus-based speech translation
Fumiaki Sugaya, Keiji Yasuda, Seiichi Yamamoto · Systems and Computers in Japan · 2004
The use of large-scale corpora in speech translation technology has not only brought us closer to a plateau of high-level speech translation capability within limited domains but also has broadened the scope of corpus gathering and helped to validate wider domain applications as well. At the same time, the efficient evaluation of speech translation systems has become an important issue. In this article, the authors propose a method of selecting a valid evaluation set based on error reduction of the translation paired comparison and discuss their experimental findings. The experimental results show how this method can be used to select a much smaller evaluation set from a large one. © 2004 Wiley Periodicals, Inc. Syst Comp Jpn, 36(1): 1–11, 2005; Published online in Wiley InterScience (www.interscience.wiley.com). DOI 10.1002/scj.20133