Evaluating MT output with entailment technology
Sebastian Padó, Michel Galley, Dan Jurafsky, Christopher D. Manning · 2008
Constant evaluation is vital to the progress of machine translation. However, human evaluation is costly, time-consuming, and difficult to do reliably. On the other hand, automatic measures of machine evaluation performance (such as BLEU, NIST, TER, and METEOR), while cheap and objective, have increasingly come under suspicion as