Non-expert evaluation of summarization systems is risky

Dan Gillick, Yang Liu · 2010

We provide evidence that intrinsic evaluation of summaries using Amazon’s Mechanical Turk is quite difficult. Experiments mirroring evaluation at the Text Analysis Conference’s summarization track show that nonexpert judges are not able to recover system rankings derived from experts. 1

Read the paper · More papers on PaperTik