Task-based MT Evaluation: From Who/When/Where Extraction to Event Understanding
Jamal Laoudi, Calandra R. Tate, Clare R. Voss · 2006
Task-based machine translation (MT) evaluation asks, how well do people perform text-handling tasks given MT output?This method of evaluation yields an extrinsic assessment of an MT engine, in terms of users' task performance on MT output.While this method is time-consuming, its key advantage is that MT users and stakeholders understand how to interpret the assessment results.Prior experiments showed that subjects can extract individual who-, when-, and where-type elements of information from MT output passages that were not especially fluent.This paper presents the results of a pilot study to assess a slightly more complex task: when given such wh-items already identified in an MT output passage, how well can subjects properly select from and place these items into wh-typed slots to complete a sentence-template about the passage's event?The results of the pilot with nearly sixty subjects, while only preliminary, indicate that this task was extremely challenging: given six test templates to complete, half of the subjects had no completely correct templates and 42% had exactly one completely correct template.The provisional interpretation of this pilot study is that event-based template completion defines a task ceiling, against which to evaluate future improvements on MT engines.