Combining Evaluation Metrics Via Loss Functions
Calandra R. Tate, Clare R. Voss · 2006
When response metrics for evaluating the utility of machine translation (MT) out-put on a given task do not yield a single ranking of MT engines, how are MT users to decide which engine best supports their task? When the cost of different types of response errors vary, how are MT users to factor that information into their rank-ings? What impact do different costs have on response-based rankings? Starting with data from an extraction ex-periment detailed in Voss & Tate (2006), this paper describes three response-rate metrics developed to quantify different as-pects of MT users ’ performance identify-ing who/when/where-items in MT output, and then presents a loss function analy-sis over these rates to derive a single cus-tomizable metric, applying a range of val-ues to correct responses and costs to dif-ferent error types. For the given experimental dataset, loss function analyses provided a clearer characterization of the engines ’ relative strength than did comparing the response rates to each other. For one MT engine, varying the costs had no impact: the en-gine consistently ranked best. By con-trast, cost variations did impact the rank-ing of the other two engines: a rank re-versal occurred on who-item extractions when incorrect responses were penalized more than non-responses. Future work with loss analysis, de-veloping operational cost ratios of error rates to correct response rates, will re-quire user studies and expert document-screening personnel to establish baseline values for effective MT engine support on wh-item extraction. 1