Task-based Evaluation of Machine Translation (MT) Engines. Measuring How Well People Extract Who, When, Where-Type Elements in MT Output

Clare R. Voss, Calandra R. Tate · 2006

How effectively can people perform the text-handling task of extracting information from the output of MT engines? When is the output of one MT engine more likely than the output of another engine to support people performing an extraction task? This paper reports on the results of a one-of-a-kind, large-scale, MT evaluation experiment where nearly sixty subjects extracted who, when, and where-type elements of information (EIs) from output generated by three types of Arabic-English MT engines. Our hypothesis was that, in an end-to-end (MT engine and user) evaluation, the best extraction results would come from subjects working with output from MT engines that reordered Arabic input to generate English word order, rather than from an engine that did not, i.e., from a statistical or a rule-based, rather than from a substitution-based MT engine. The results of the experiment were not as straight-forward as expected: (1) non-response rates were statistically comparable across all three evaluated MT engines, while (2) correct response rates were statistically comparable on two engines, the statistical and substitution-based engines that yielded better (higher) rates than the rule-based engine did, and (3) incorrect response rates were statistically comparable on a different pair of two engines, the rule- and substitution-based engines that yielded significantly worse (higher) rates than the statistical engine. While these results do indicate that the statistical engine yielded significantly better rates than at least one of the other two engines on two of the three metrics, the lack of uniform results pre-empts an across-the-board ranking of the engines. Our next step is to incorporate the collected data in statistical models and test for their adequacy in predicting these task results from faster and less expensive, automatic metrics. The long-term goal is to understand which metrics accurately predict MT users ’ task effectiveness with different MT engines on text-handling tasks of varying levels of difficulty. 1

Read the paper · More papers on PaperTik