Simulating Human Judgment in Machine Translation Evaluation Campaigns
Philipp Koehn · Edinburgh Research Explorer (University of Edinburgh) · 2012
We present a Monte Carlo model to simulate human judgments in machine translation evaluation campaigns, such as WMT or IWSLT. We use the model to compare different ranking methods and to give guidance on the number of judgments that need to be collected to obtain sufficiently significant distinctions between systems