Non-expert evaluation of summarization systems is risky
Dan Gillick, Yang Liu · 2010
We provide evidence that intrinsic evaluation of summaries using Amazon’s Mechanical Turk is quite difficult. Experiments mirroring evaluation at the Text Analysis Conference’s summarization track show that nonexpert judges are not able to recover system rankings derived from experts. 1