How Biased Is Your NLG Evaluation?
Pavlos Vougiouklis, Eddy Maddalena, Jonathon S. Hare, Elena Paslaru Bontas Simperl · ePrints Soton (University of Southampton) · 2018
Human assessments by either experts or crowdworkers are used extensively for the evaluation of systems employed on a variety of text generative tasks. In this paper, we focus on the human evaluation of textual summaries from knowledge base triple-facts. More specifically, we investigate possible similarities between the evaluation that is performed by experts and crowdworkers. We generate a set of summaries from DBpedia triples using a state-of-the-art neural network architecture. These summaries are evaluated against a set of criteria by both experts and crowdworkers. Our results highlight significant differences between the scores that are provided by the two groups.