Multi-level Evaluation for Machine Translation
Boxing Chen, Hongyu Guo, Roland Kühn · 2015
Translations generated by current statistical systems often have a large variance, in terms of their quality against human references.To cope with such variation, we propose to evaluate translations using a multi-level framework.The method varies the evaluation criteria based on the clusters to which a translation belongs.Our experiments on the WMT metric task data show that the multi-level framework consistently improves the performance of two benchmarking metrics, resulting in better correlation with human judgment.