A Re-examination of Machine Learning Approaches for Sentence-Level MT Evaluation
Joshua S. Albrecht, Rebecca Hwa · 2007
Machine learning methods have been proposed in the past as means of developing automatic metrics to evaluate the quality of machine translated sentences. This paper further investigates this idea, analyzing aspects of learning that impact performance. We show that previously proposed approaches of training a Human-Likeness classifier is not as well correlated with human judgments of translation quality. Instead, we argue that regression-based learning produces more reliable metrics. We demonstrate the feasibility of regressionbased metrics through empirical analysis of learning curves and generalization studies. Our results suggest that regression-based metrics can achieve higher correlations with human judgments than several standard automatic metrics. 1